Filesystem MCP Server
README.md
# MCP Integration — Resume Matching Agent
A production-ready implementation of the **Model Context Protocol (MCP)** applied to an AI-powered resume matching system. The project demonstrates how to replace custom file-system tools with a standardised MCP server and connect a LangGraph agent to it via the MCP client.
---
## Table of Contents
1. [Overview](#overview)
2. [Architecture](#architecture)
3. [Project Structure](#project-structure)
4. [Setup](#setup)
5. [Usage](#usage)
6. [MCP Server Reference](#mcp-server-reference)
7. [Agent Workflow](#agent-workflow)
8. [Test Scenarios](#test-scenarios)
9. [Sample Output](#sample-output)
---
## Overview
| Layer | Technology |
|-------|-----------|
| MCP Server | Python `mcp` SDK (FastMCP), stdio transport, JSON-RPC 2.0 |
| Agent Framework | LangGraph `StateGraph` (explicit state machine) |
| MCP Client | `langchain-mcp-adapters` `MultiServerMCPClient` |
| LLM | Anthropic Claude (`claude-sonnet-5`) via `langchain-anthropic` |
| Concurrency | `ThreadPoolExecutor` inside `batch_process` |
The agent never touches the filesystem directly — every read, write, and directory operation is a JSON-RPC 2.0 call to the MCP server subprocess.
---
## Architecture
```
┌──────────────────────────────────────────────────────────────┐
│ matching_agent.py │
│ │
│ LangGraph StateGraph │
│ load_jd → extract_req → fetch_resumes │
│ → analyze → rank → report → save │
│ ↘ error_handler ↙ │
│ │
│ MultiServerMCPClient (langchain-mcp-adapters) │
└──────────────────────────┬───────────────────────────────────┘
│ stdio (JSON-RPC 2.0)
┌──────────────────────────▼───────────────────────────────────┐
│ filesystem_mcp_server.py │
│ │
│ FastMCP — 9 tools + 2 resources │
│ Milestone-1 : read_file write_file list_directory │
│ search_files get_file_info delete_file │
│ copy_file │
│ MCP-specific: watch_directory batch_process │
│ Resources : config://server filesystem://resumes │
└──────────────────────────────────────────────────────────────┘
│
Local Filesystem
data/resumes/ data/job_descriptions/ data/results/
```
---
## Project Structure
```
MCPIntegration/
├── filesystem_mcp_server.py # MCP server (JSON-RPC 2.0, stdio)
├── skills_db_mcp_server.py # 2nd MCP server — labour-market DB (multi-MCP bonus)
├── matching_agent.py # LangGraph agent with multi-MCP client
├── run_tests.py # 13 test scenarios
├── requirements.txt # Python dependencies
├── workflow_diagram.md # State machine & protocol diagrams
└── data/
├── resumes/
│ ├── alice_chen.txt # Senior ML Engineer (strong match)
│ ├── bob_martinez.txt # Full-stack dev (partial match)
│ └── carol_johnson.txt # Data Scientist / ML Eng (good match)
├── job_descriptions/
│ └── senior_ml_engineer.txt
└── results/ # Generated reports land here
```
---
## Setup
### Prerequisites
- Python 3.10 or later
- An Anthropic API key (only needed for the agent; tests run without it)
### Install dependencies
```bash
pip install -r requirements.txt
```
### Set API key
```bash
# macOS / Linux
export ANTHROPIC_API_KEY=sk-ant-...
# Windows (PowerShell)
$env:ANTHROPIC_API_KEY = "sk-ant-..."
# Windows (Command Prompt)
set ANTHROPIC_API_KEY=sk-ant-...
```
---
## Usage
### Run the MCP server standalone (inspect mode)
```bash
python -m mcp dev filesystem_mcp_server.py
```
### Run the full resume matching agent
```bash
python matching_agent.py \
--job data/job_descriptions/senior_ml_engineer.txt \
--resumes data/resumes \
--output data/results
```
Options:
| Flag | Default | Description |
|------|---------|-------------|
| `--job` | `data/job_descriptions/senior_ml_engineer.txt` | Path to job description file |
| `--resumes` | `data/resumes` | Directory of candidate `.txt` files |
| `--output` | `data/results` | Output directory for report and scores |
| `--model` | `claude-sonnet-5` | Anthropic model ID |
The agent writes two files to the output directory on completion:
- `match_report_<timestamp>.md` — executive Markdown report
- `scores_<timestamp>.json` — structured per-candidate scores
### Run test scenarios (no API key required)
```bash
python run_tests.py
```
### Run tests including the end-to-end agent
```bash
python run_tests.py --e2e
```
---
## MCP Server Reference
All tools return a JSON object with a `"status"` field (`"success"` or `"error"`).
Error strings are prefixed with an error code, e.g. `"FILE_NOT_FOUND: ./x.txt"`.
### Milestone-1 Tools
#### `read_file(path)`
Read the text content of a file.
```json
{ "status": "success", "path": "...", "content": "...", "size": "4.2 KB" }
```
#### `write_file(path, content, overwrite=true)`
Write text to a file; creates parent directories automatically.
```json
{ "status": "success", "path": "...", "bytes_written": 1234, "size": "1.2 KB" }
```
#### `list_directory(path=".", pattern="*", recursive=false)`
List files and sub-directories with optional glob filtering.
```json
{
"status": "success",
"count": 3,
"entries": [
{ "name": "alice_chen.txt", "type": "file", "size": "2.1 KB", "modified": "..." }
]
}
```
#### `search_files(directory, query, file_extensions=".txt,.md,.pdf")`
Case-insensitive full-text search. Returns up to 10 matching lines per file.
```json
{ "status": "success", "files_matched": 2, "results": [ { "filename": "...", "matches": [...] } ] }
```
#### `get_file_info(path)`
Rich metadata including MD5 checksum (files) or child counts (directories).
```json
{ "status": "success", "name": "alice_chen.txt", "size": "2.1 KB", "md5_checksum": "a1b2c3..." }
```
#### `delete_file(path)`
Remove a file (not a directory).
#### `copy_file(source, destination)`
Copy a file with metadata; creates destination parent dirs.
---
### MCP-Specific Capabilities
#### `watch_directory(path, duration_seconds=30, file_extensions=".txt,.pdf,.docx,.md")`
Polls a directory for change events during the specified window (max 300 s).
Returns a list of `created`, `modified`, and `deleted` events.
```json
{
"status": "success",
"events_detected": 2,
"events": [
{ "type": "created", "filename": "new_resume.txt", "elapsed_seconds": 4.1 }
]
}
```
**Use case:** detect newly uploaded resumes without restarting the server.
#### `batch_process(directory, operation, file_pattern="*.txt", max_workers=4)`
Processes all matching files concurrently using a thread pool (1–8 workers).
| `operation` | Output per file |
|-------------|-----------------|
| `read_all` | Full text content |
| `index` | Word count, line count, char count, size, modified date |
| `extract_skills` | List of detected technical skill keywords |
| `summarize` | First 5 lines + word count + top 10 skills |
```json
{
"status": "success",
"processed_count": 3,
"elapsed_seconds": 0.012,
"results": [ { "file": "alice_chen.txt", "skill_count": 22, "skills": ["python", ...] } ]
}
```
---
### MCP Resources
Resources are discoverable via `resources/list` and readable via `resources/read`.
| URI | Description |
|-----|-------------|
| `config://server` | Live server configuration (tools list, size limits, supported extensions) |
| `filesystem://resumes` | Index of all resume files in the configured resume directory |
---
## Agent Workflow
The agent is a six-node LangGraph `StateGraph`. Nodes in **bold** make LLM calls; nodes in *italics* call MCP tools.
```
START
│
▼
[1] load_job_description ← MCP: read_file
│
▼
[2] extract_requirements ← LLM: parse JD into structured dict
│
▼
[3] fetch_resumes ← MCP: list_directory + batch_process + read_file × N
│
▼
[4] analyze_matches ← LLM: score each resume 0–100 against requirements
│
▼
[5] rank_candidates ← Python: sort by overall_score descending
│
▼
[6] generate_report ← LLM: write executive Markdown report
│
▼
[7] save_results ← MCP: write_file × 2 (report + scores JSON)
│
▼
END
Any node failure → error_handler → END
```
### State object (key fields)
| Field | Populated by | Type |
|-------|-------------|------|
| `job_description` | `load_job_description` | `str` |
| `job_requirements` | `extract_requirements` | `dict` |
| `resume_contents` | `fetch_resumes` | `dict[str, str]` |
| `match_scores` | `analyze_matches` | `list[dict]` |
| `ranked_candidates` | `rank_candidates` | `list[dict]` |
| `final_report` | `generate_report` | `str` |
| `report_path` | `save_results` | `str` |
---
## Test Scenarios
`run_tests.py` covers 12 independent test groups against the live MCP server:
| # | Test | What it checks |
|---|------|---------------|
| 1 | Server connectivity | All 9 tools discovered via `tools/list` |
| 2 | `read_file` | Success path + `FILE_NOT_FOUND` error |
| 3 | `write_file` | Write, read-back verify, `overwrite=False` error |
| 4 | `list_directory` | Count ≥ 3 resumes, recursive flag |
| 5 | `search_files` | Keyword hit across ≥ 2 files, zero-result case |
| 6 | `get_file_info` | MD5 checksum present, directory child counts |
| 7 | `batch_process / index` | Word/line counts for all resumes |
| 8 | `batch_process / extract_skills` | Skills list per resume |
| 9 | `batch_process / summarize` | First-5-lines preview |
| 10 | `watch_directory` | Detects a file created mid-window |
| 11 | `copy_file` | Copy verified via `get_file_info` |
| 12 | `delete_file` | Removes temp files from tests 3 and 11 |
| E2E | Full agent run | End-to-end with real LLM (requires API key) |
---
## Sample Output
```
════════════════════════════════════════════════════════════
RESUME MATCHING AGENT · MCP + LangGraph + Claude
════════════════════════════════════════════════════════════
MCP tools available: ['batch_process', 'copy_file', 'delete_file',
'get_file_info', 'list_directory', 'read_file', 'search_files',
'watch_directory', 'write_file']
[1/6] Loading job description…
2,134 characters loaded
[2/6] Extracting structured requirements with LLM…
Position : Senior Machine Learning Engineer
Required : 8 skills
Preferred : 6 skills
[3/6] Fetching resumes from 'data/resumes'…
Found 3 resume file(s). Batch-indexing…
✓ alice_chen.txt (412 words, 63 lines)
✓ bob_martinez.txt (287 words, 54 lines)
✓ carol_johnson.txt (351 words, 61 lines)
[4/6] Analysing 3 resume(s)…
Scoring alice_chen.txt… 94/100 — Strong Match
Scoring carol_johnson.txt… 81/100 — Good Match
Scoring bob_martinez.txt… 38/100 — Partial Match
[5/6] Ranking candidates…
#1 Alice Chen 94/100 Strong Match
#2 Carol Johnson 81/100 Good Match
#3 Bob Martinez 38/100 Partial Match
[6/6] Generating final report…
Report generated (3,847 characters).
Report → data/results/match_report_20260624_143022.md
Scores → data/results/scores_20260624_143022.json
```
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues