Skip to main content
Glama
jstrick9

MCP Web Research Agent

by jstrick9
README.md
# MCP Web Research Agent for macOS

A Python [Model Context Protocol](https://modelcontextprotocol.io) (MCP) agent/server that gives a local AI assistant tools for:

- `search_web` — public web search via DuckDuckGo HTML results
- `fetch_url` — fetch a public web page and extract readable text
- `save_note` — save research notes as Markdown/text files in a folder you choose

The project also includes `agent.py`, a small local bridge that connects **Ollama** to the MCP server. Ollama runs the LLM; this project provides the MCP tools and the tool-calling agent loop.

> Note: Ollama itself is a model server, not a native MCP client. To use Ollama with MCP tools, run `agent.py` here or another MCP bridge/client.

## What you need

- MacBook Pro with macOS
- Python 3.11 or newer (the `mcp` package requires Python 3.10+; setup prefers 3.13/3.12/3.11)
- [Ollama](https://ollama.com) installed and running
- A tool-calling local model. Recommended starting point:
  - `qwen2.5:7b` for 16 GB RAM Macs
  - `qwen2.5:14b` if you have enough RAM/performance
  - `qwen3:14b` if your Ollama version supports it well

## 1. Install

Open Terminal and run:

```bash
mkdir -p ~/Agents
cd ~/Agents
git clone https://github.com/jstrick9/mcp_agent.git
cd mcp_agent

python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r requirements.txt
```

`setup-mac.sh` does the same thing and will find or install Python 3.11+ for you:

```bash
bash setup-mac.sh
```

If you do not have Python 3.11+:

```bash
brew install python
```

### Confirm the install works (no Ollama needed)

These two checks start the real MCP servers and drive the real agent loop
against a mock Ollama endpoint. They need no network and no downloaded model,
so they are the fastest way to confirm a fresh clone is healthy:

```bash
./.venv/bin/python tests/e2e_mcp.py
bash tests/e2e_agents.sh
```

You should see `ALL CHECKS PASSED` and `ALL BRIDGE AGENT CHECKS PASSED`.
Together they exercise all 43 MCP tools plus one full tool call through each
bridge agent.

## 2. Install and start Ollama

Install Ollama from <https://ollama.com> or with Homebrew:

```bash
brew install --cask ollama
```

Open the Ollama app once, then pull a model:

```bash
ollama pull qwen2.5:7b
ollama serve
```

In another Terminal tab, verify Ollama is running:

```bash
curl http://localhost:11434/api/tags
```

## 3. Run the local Ollama MCP agent

From the project folder:

```bash
cd ~/Agents/mcp_agent
source .venv/bin/activate
python agent.py
```

Then ask something like:

```text
Research recent MCP news, open the two best sources, summarize them, and save the summary as mcp-news.md.
```

One-shot mode:

```bash
python agent.py "Research current MCP SDK best practices and save notes."
```

Use a different model:

```bash
python agent.py --model qwen2.5:14b
```

Choose where notes are saved:

```bash
python agent.py --notes-dir ~/Documents/research-notes
```

## 4. Use with Claude Desktop

### All five servers at once

To register every agent in this repo with one paste, start from the generated combined config:

- `claude_desktop_config.all.example.json` — all five servers for Claude Desktop
- `cursor-mcp.all.example.json` — all five servers for Cursor

Replace `YOUR_USERNAME` with your macOS username, then merge the `mcpServers` object into:

```text
~/Library/Application Support/Claude/claude_desktop_config.json
```

Both files declare all five servers — `web-research`, `local-planner`, `health-tracker`, `knowledge-base`, and `flashcards` — each pointing at the shared `.venv/bin/python` and its own data directory.

These two files are **generated**, not hand-written. They are assembled from the per-agent example configs by:

```bash
./.venv/bin/python tools/build_combined_configs.py          # regenerate
./.venv/bin/python tools/build_combined_configs.py --check   # verify they are current
```

`--check` also confirms each entry's environment variable matches what the server script actually reads, so a renamed `MCP_*` variable cannot silently ship a broken config. The check runs as part of `tests/e2e_mcp.py`.

### One server at a time

If you want Claude Desktop to connect directly to the MCP server, edit:

```text
~/Library/Application Support/Claude/claude_desktop_config.json
```

Add:

```json
{
  "mcpServers": {
    "web-research": {
      "command": "/Users/YOUR_USERNAME/Agents/mcp_agent/.venv/bin/python",
      "args": [
        "/Users/YOUR_USERNAME/Agents/mcp_agent/server.py"
      ],
      "env": {
        "MCP_NOTES_DIR": "/Users/YOUR_USERNAME/MCPWebResearch/notes"
      }
    }
  }
}
```

Replace `YOUR_USERNAME` with your Mac username. Create the file if it does not exist. Restart Claude Desktop after editing.

## 5. Use with Cursor

Create or edit `.cursor/mcp.json` in a workspace:

```json
{
  "mcpServers": {
    "web-research": {
      "command": "/Users/YOUR_USERNAME/Agents/mcp_agent/.venv/bin/python",
      "args": [
        "/Users/YOUR_USERNAME/Agents/mcp_agent/server.py"
      ],
      "env": {
        "MCP_NOTES_DIR": "/Users/YOUR_USERNAME/MCPWebResearch/notes"
      }
    }
  }
}
```

Then restart Cursor or reload its MCP settings.

## Tool reference

### `search_web(query: str, max_results: int = 5)`

Returns search results as JSON:

```json
{
  "query": "Model Context Protocol",
  "results": [
    {
      "title": "Example",
      "url": "https://example.com",
      "snippet": "..."
    }
  ]
}
```

### `fetch_url(url: str, max_chars: int = 8000)`

Fetches an `http`/`https` URL and returns extracted text. It avoids JavaScript rendering, so it works best on normal HTML pages.

### `save_note(filename: str, content: str)`

Saves a note to `MCP_NOTES_DIR`. The default directory is:

```text
~/MCPWebResearch/notes
```

The tool sanitizes filenames and blocks path traversal.

## Configuration

Environment variables:

- `OLLAMA_MODEL` — default model used by `agent.py`; default is `qwen2.5:7b`
- `OLLAMA_URL` — OpenAI-compatible Ollama chat endpoint; default is `http://localhost:11434/v1/chat/completions`
- `MCP_NOTES_DIR` — directory for saved notes

Example:

```bash
export OLLAMA_MODEL=qwen2.5:14b
export MCP_NOTES_DIR=~/Documents/research-notes
python agent.py
```

## Troubleshooting

### `Connection refused` to `localhost:11434`

Ollama is not running. Start it with:

```bash
ollama serve
```

### The agent does not call tools

Use a model with strong tool-calling support. `qwen2.5:7b`, `qwen2.5:14b`, and similar Qwen models are good starting points.

### A page returns little text

Some websites block non-browser clients or require JavaScript. Try a different source, or use `search_web` and `fetch_url` together.

### Claude Desktop does not show the server

Double-check that:

- The Python path points to `.venv/bin/python` inside this project
- The `server.py` path is absolute
- The JSON file has valid syntax
- You fully restarted Claude Desktop

## Files

- `server.py` — MCP server with web research tools
- `agent.py` — local Ollama-powered MCP client/agent loop
- `requirements.txt` — Python dependencies

## Safety notes

- This server can fetch public URLs and search the public web.
- It can write files only into `MCP_NOTES_DIR`.
- It does not execute shell commands.
- Review saved notes and citations before relying on them.

---

# Second MCP agent: Local Planner

The repo now includes a second MCP server/agent: `local-planner`. It stores projects, tasks, and Markdown notes on disk.

## What it does

Tools:

- `create_project(name, description)`
- `list_projects()`
- `create_task(project, title, notes, priority, due_date, status)`
- `list_tasks(project, status)`
- `update_task(project, task_id, ...)`
- `complete_task(project, task_id)`
- `delete_task(project, task_id)`
- `save_project_note(project, content, append)`
- `get_daily_focus(for_date)`

Default data directory:

```text
~/MCPPlanner
```

Override it with:

```bash
export MCP_PLANNER_DIR=~/Documents/my-planner
```

## Run the Ollama planner

```bash
bash run-planner.sh
```

One-shot:

```bash
bash run-planner.sh "Create a project called Weekend Yard Work with tasks for mowing, trimming bushes, and buying mulch."
```

Use a different model:

```bash
bash run-planner.sh --model qwen2.5:14b
```

Store planner data elsewhere:

```bash
bash run-planner.sh --data-dir ~/Documents/planner-data
```

## Good planner prompts

```text
Create a project called Home Network Upgrade and break it into at least six tasks with priorities.
```

```text
Look at my daily focus and tell me what I should work on first.
```

```text
Create a moving checklist project with tasks, due dates, and notes.
```

```text
Mark the first task in the Weekend Yard Work project complete and tell me what remains.
```

## Claude Desktop config for planner

Use:

```text
claude_desktop_config.planner.example.json
```

Add it to:

```text
~/Library/Application Support/Claude/claude_desktop_config.json
```

You can merge both servers under `mcpServers` so Claude sees web research and planning tools.

## Cursor config for planner

Use:

```text
cursor-mcp.planner.example.json
```

## Files

- `planner_server.py` — MCP server for projects/tasks/notes
- `planner_agent.py` — Ollama bridge/agent for the planner
- `run-planner.sh` — launcher

---

# Third MCP agent: Health & Habit Tracker

The repo includes a third MCP server/agent for tracking habits, workouts, meals, and body measurements. All data is stored locally under `MCP_HEALTH_DIR` (default `~/MCPHealth`).

> This tool is for personal tracking only and does not provide medical advice.

## Tools

- `create_habit(name, description, target_per_week, unit)`
- `list_habits(active_only)`
- `log_habit(log_date, value, notes, habit_id|habit_name)`
- `log_workout(activity, duration_minutes, log_date, intensity, calories, distance_km, notes)`
- `log_meal(description, meal_type, log_date, calories, protein_g, carbs_g, fat_g, notes)`
- `log_measurement(weight_kg, log_date, body_fat_pct, waist_cm, notes)`
- `list_logs(log_type, from_date, to_date, limit)`
- `delete_log(log_id)`
- `save_health_note(content, append)`
- `get_daily_summary(for_date)`
- `get_weekly_report(for_date)`

## Run with Ollama

```bash
bash run-health.sh
```

One-shot:

```bash
bash run-health.sh "Create habits for a 30-min walk, drinking water, and stretching, then log today's walk and a lunch salad."
```

Use a different model or data directory:

```bash
bash run-health.sh --model qwen2.5:14b --data-dir ~/Documents/health-data
```

## Good prompts

```text
Create habits for walking 5 days per week, drinking 80 oz of water, and stretching daily.
```

```text
Log a 45-minute moderate run today that burned 420 calories and covered 6 km.
```

```text
Log my breakfast: oatmeal with banana and peanut butter, about 520 calories and 22 grams of protein.
```

```text
Give me today's health summary and list habits I still need to complete.
```

```text
Give me my weekly report and tell me which habits I'm behind on.
```

## Claude Desktop / Cursor configs

- `claude_desktop_config.health.example.json`
- `cursor-mcp.health.example.json`

You can run all five MCP servers together (web research, planner, health, knowledge base, flashcards) by listing each under `mcpServers`.

## Files

- `health_server.py` — MCP server
- `health_agent.py` — Ollama bridge/agent
- `run-health.sh` — launcher

---

# Fourth MCP agent: Personal Knowledge Base

A searchable long-term memory that ties the other three agents together. Your research agent, planner, and health tracker all *write* notes, but nothing could search them. This agent indexes those folders plus anything you save manually, with real full-text search.

Search uses SQLite's FTS5 extension with BM25 relevance ranking (both ship with Python, so there are no new dependencies). If FTS5 is unavailable on a platform, the server automatically falls back to substring search and reports `"search_mode": "substring"`.

Data is stored under `MCP_KB_DIR` (default `~/MCPKnowledge`) in a single SQLite file, `kb.db`.

## Tools

- `save_snippet(content, title, tags, source_url, source_path, source_type)`
- `search_kb(query, tag, limit, content_chars)`
- `list_snippets(tag, source_type, limit, content_chars)`
- `get_snippet(snippet_id)`
- `delete_snippet(snippet_id)`
- `list_tags()`
- `rename_tag(old_tag, new_tag)`
- `ingest_notes(directories, tag, recursive)`
- `kb_stats()`

## Import notes from your other agents

This is the highest-value first step. It pulls `.md` and `.txt` files into the index:

```text
Ingest my notes from ~/MCPWebResearch/notes, ~/MCPPlanner, and ~/MCPHealth with the tag imported.
```

Re-running is safe: unchanged files are skipped, changed files are updated in place, and nothing is duplicated. Skips `.git`, `.venv`, and `node_modules`.

## Run with Ollama

```bash
bash run-kb.sh
```

One-shot:

```bash
bash run-kb.sh "Ingest my research notes, then summarize everything I have saved about MCP."
```

Different model or database location:

```bash
bash run-kb.sh --model qwen2.5:14b --data-dir ~/Documents/knowledge
```

## Good prompts

```text
Save this: FastMCP exposes Python functions as MCP tools via a decorator. Tag it python and mcp.
```

```text
Search my knowledge base for stdio transport and cite the source of each result.
```

```text
What do I have tagged mcp? List titles and sources.
```

```text
Rename the tag py to python everywhere.
```

```text
Give me knowledge base stats: how many entries, which sources, and my top tags.
```

## Search syntax

`search_kb` passes your query to SQLite FTS5, so operators work:

| Query | Meaning |
|---|---|
| `stdio transport` | entries containing both words |
| `"stdio transport"` | that exact phrase |
| `mcp OR sourdough` | either word |
| `mcp NOT health` | `mcp` but not `health` |
| `protocol*` | prefix match |

If a query contains malformed FTS5 syntax, the server retries it as literal quoted phrases rather than failing, so unusual punctuation never produces an error.

## Claude Desktop / Cursor configs

- `claude_desktop_config.kb.example.json`
- `cursor-mcp.kb.example.json`

You can run all four MCP servers together by listing each under `mcpServers`.

## Files

- `kb_server.py` — MCP server with SQLite FTS5 search
- `kb_agent.py` — Ollama bridge/agent
- `run-kb.sh` — launcher

## Safety notes

- The database and all writes stay inside `MCP_KB_DIR`.
- `ingest_notes` reads only the folders you explicitly pass to it.
- It reads `.md` and `.txt` files only, and never executes shell commands.
- Notes may contain personal information. Back up `kb.db` like any other data file.

---

# Fifth MCP agent: Learning & Flashcards

A spaced-repetition study system. Build decks from anything you're learning, then review cards on a schedule that adapts to how well you actually remember them.

Scheduling uses the **SM-2 algorithm** (SuperMemo 2), the same family of algorithms Anki is built on. Cards you find hard come back sooner; cards you know well stretch out to weeks and months.

Data is stored under `MCP_FLASHCARDS_DIR` (default `~/MCPFlashcards`) as three JSON files: `decks.json`, `cards.json`, and `reviews.json`.

## Tools

- `create_deck(name, description, tags)`
- `list_decks()`
- `delete_deck(deck)`
- `add_card(deck, front, back, tags, notes)`
- `edit_card(card_id, front, back, tags, notes)`
- `delete_card(card_id)`
- `list_cards(deck, tag, limit, due_only)`
- `get_due_cards(deck, limit, include_back)`
- `record_review(card_id, quality, seconds_spent, session_id, notes)`
- `get_review_session(session_id, for_date)`
- `get_stats(deck)`

## How grading works

`record_review` takes a `quality` score from 0 to 5:

| Quality | Meaning | Effect |
|---|---|---|
| 0 | Complete blackout | Lapse — card restarts, due tomorrow |
| 1–2 | Wrong, or recalled only after seeing the answer | Lapse — card restarts |
| 3 | Correct, but with serious difficulty | Passes, ease factor drops |
| 4 | Correct after hesitation | Passes, ease factor holds steady |
| 5 | Perfect recall | Passes, ease factor rises |

Intervals progress 1 day, then 6 days, then each interval is the previous one multiplied by the card's ease factor (which starts at 2.5 and is clamped between 1.3 and whatever repeated perfect reviews build it to).

## Run with Ollama

```bash
bash run-flashcards.sh
```

One-shot:

```bash
bash run-flashcards.sh "Create a deck called MCP Basics with 8 cards covering the protocol fundamentals."
```

Different model or data directory:

```bash
bash run-flashcards.sh --model qwen2.5:14b --data-dir ~/Documents/flashcards
```

Reviews are interactive, so the REPL is the better fit:

```bash
bash run-flashcards.sh
```

```text
You> What's due today? Quiz me on it.
```

The agent will show you each question, wait for your answer, then ask you to rate your recall from 0 to 5. It does not grade itself.

## Good prompts

```text
Create a deck called Spanish Food with 10 cards for common restaurant vocabulary.
```

```text
Make flashcards from this: MCP servers expose tools over stdio, use JSON-RPC 2.0, and are spawned by the client.
```

```text
Quiz me on MCP Basics, ten cards.
```

```text
What's my retention this week and which cards do I keep failing?
```

```text
Show my study stats and how many cards are due tomorrow.
```

## Pair it with the knowledge base

The knowledge base agent can pull in notes you've saved, which makes good raw material for cards:

```text
Search my knowledge base for mcp and make a flashcard deck from what you find.
```

## Claude Desktop / Cursor configs

- `claude_desktop_config.flashcards.example.json`
- `cursor-mcp.flashcards.example.json`

You can run all five MCP servers together — see `claude_desktop_config.all.example.json` in the [Use with Claude Desktop](#4-use-with-claude-desktop) section.

## Files

- `flashcards_server.py` — MCP server with SM-2 scheduling
- `flashcards_agent.py` — Ollama bridge/agent
- `run-flashcards.sh` — launcher

## Safety notes

- All data stays inside `MCP_FLASHCARDS_DIR`.
- The server never executes shell commands or makes network calls.
- Review history is retained when you delete a deck, so `get_stats` stays meaningful.