Job Application Tracker
by hanjot
README.md
# Job Application Tracker (MCP Server)


**[Live dashboard →](https://hanjot.github.io/job-application-tracker-mcp/)**
A real [Model Context Protocol](https://modelcontextprotocol.io) server that
turns my own job search — 200 real applications sent between August 10 and
September 29, 2026 — into something an AI assistant can query and update in
plain language.
This is a follow-up to my first MCP project
([mcp-filesystem-connection](https://github.com/hanjot/mcp-filesystem-connection)),
which proved the basic AI-to-local-file connection. This one is built on my
actual job-search data and does real read/write work: querying, filtering,
and updating applications through MCP tools, not just reading a file.
## Why I built it — and what it's actually for
I'm a Lead Technical Program Manager currently in an active job search, and
by the time I built this I had sent 150+ applications with no single place
to see them — just a folder of PDFs named inconsistently
(`Company_Role_Date.pdf`, `RoleCompanyDate.pdf`, some with location, some
without).
This tracker is deliberately **not** about callbacks or interview status.
The question I actually care about is: across everything I've applied to,
what skills is the market asking for — Salesforce, SAP, cybersecurity, AI,
and to what depth — and what is each of those actually paying? Seeing that
pattern across 200 applications, instead of one job description at a time,
is the actual value.
## What it does
`parse_applications.py` reads the raw filenames and file timestamps from my
Applications folder and parses each one into: company, role, applied date,
a `skills_required` tag list (Salesforce, SAP, AI/ML, AI/GenAI,
Cybersecurity, Cloud, Data/Analytics, Compliance, Payments, Agile, Data
Privacy, and so on), and — where a pay range is either stated in a saved
posting or known from a real conversation about that specific posting — a
`salary_range`.
`server.py` is the MCP server. It loads `applications.json` (the parsed
output) and exposes these tools to any MCP client:
| Tool | What it does |
|---|---|
| `list_applications` | Filter by company, by required skill, or by date |
| `get_application` | Full detail on one application by id |
| `add_application` | Log a new application, with the real job link + job description text if available (auto-tags skills and extracts a salary range from the text when present) |
| `update_application` | Correct company/role/notes, add a real `job_url`/`job_description` retroactively (auto re-tags `skills_required` and re-extracts `salary_range` from the real text), or set skills/salary by hand |
| `get_summary` | Totals, date range, repeat companies, and — the main point — **skills_frequency**: how often each skill/technology showed up across every posting |
| `get_salary_summary` | Average disclosed pay by skill category (e.g. AI vs. Cybersecurity vs. Salesforce) — hourly rates annualized so everything compares on the same basis |
| `get_course_priority` | Ranks which course/certification to prioritize next, based on real demand across all 200 applications (pulls from my existing AI Course Priority Plan where a matching course exists) |
| `find_duplicates` | Companies applied to more than once, with each application listed |
`skills_report.html` is a standalone chart + table view of the same data
(open it in any browser) — a horizontal bar chart of skill frequency across
all 200 applications, a second chart of average disclosed salary by skill
category, and the course-priority ranking below both.
**Real job descriptions, going forward:** each application also has a
`job_url` and `job_description` field. When a real posting's text is saved
(via `add_application` or `update_application`), `skills_required` is
automatically re-tagged from that actual text instead of guessed from the
title, and `salary_range` is extracted automatically if the text states a
pay range — the same keyword matcher, just run against real content. This
is how new applications get added from here on: job link + full JD text in,
accurate skill tags and pay data out.
## Real numbers from my own search (as of Sept 29, 2026)
- **200 applications** tracked, spanning **Aug 10 – Sept 29, 2026**, across
**133 unique companies**
- Skills flagged from job titles (and real posting text where saved) so
far: **Program/Project Management** (135), **AI/GenAI/ML** (22),
**Cybersecurity** (11), **Data/Analytics** (9),
**Compliance/Risk/Governance** (6), **Infrastructure/DevOps** (5), plus
smaller counts for Agile/Scrum, Cloud, Payments, Salesforce, and Data
Privacy
- 48 applications have no skill tag yet — their titles didn't contain a
recognizable keyword (see limitation below)
- **35 of 200 applications have a disclosed salary range** — 2 from a real
saved job description, the rest from pay ranges discussed for that
specific posting. By skill category, average annualized base pay ranges
from roughly **$190K (AI/GenAI/ML)** to **$243K (Cybersecurity)** among
the categories with disclosed data; see `get_salary_summary` or the
dashboard for the full breakdown. Categories with zero disclosed data
simply don't appear — nothing here is a guessed or averaged-out figure.
## Running it
```bash
pip install -r requirements.txt
python parse_applications.py # rebuilds applications.json from raw_listing.json
mcp dev server.py # opens the MCP Inspector to try the tools by hand
```
To connect it to an MCP-compatible client, add it as a stdio server that
runs `python server.py` from this folder.
## Verifying it's real
Two layers of testing, both run automatically on every push via GitHub
Actions (see the CI badge above):
1. **Unit tests** (`tests/test_parse_applications.py`, run with `pytest`) —
check the filename-parsing logic against real, tricky cases from the
actual dataset: camelCase titles with no separators, a company name
("Marketing") that contains a month abbreviation as a substring
("mar"), dates that must fall back to the file's save time when the
filename has none, self-authored resume filenames that start with my
own name, flagged duplicate files, and salary-range text extraction.
2. **End-to-end protocol test** (`test_client.py`) — spawns `server.py` as
an actual MCP server over stdio, calls `initialize`, lists the tools,
and calls `get_summary`, `get_salary_summary`, `list_applications`,
`add_application`, `update_application`, `get_course_priority`, and
`find_duplicates` — reading back real results and confirming the write
tools actually persisted changes to `applications.json`. That's a
genuine protocol round trip, not a mocked call.
Run both locally with:
```bash
pip install -r requirements.txt pytest
python parse_applications.py
python -m pytest tests/ -v
python test_client.py
```
## Honest limitations
- **For most of the 200 applications parsed from filenames,
`skills_required` is still inferred from the job title only** — I don't
have the original posting text saved for those, so it's a best-effort
keyword match against the role text, not a transcription of each
posting's actual requirements. Going forward, any application added with
a real `job_description` gets accurate, text-based tags instead (see
above); I'm backfilling the older ones with `update_application` as I
revisit real postings.
- **Salary data is partial and curated, not exhaustive.** Only 35 of 200
applications have a disclosed range — either extracted from a saved job
description, or entered by hand from a specific pay conversation I had
about that posting. I deliberately did not guess or average a figure for
postings without a real disclosed number, so most applications (and a
few whole skill categories) simply have no salary data yet.
- The filename parser is heuristic in general. A small number of
applications (roughly 5%) come through with an unclear role or company
because the original filename ran words together with no separator.
- There is no status/callback field by design — this tracker is about
required skills (and pay) across the whole search, not per-application
outcomes.
- Four files that were saved copies of my base resume, or duplicate
`.docx` copies of a `.pdf` I'd already counted, are excluded from the
counts rather than counted as applications.
## Stack
Python, the official [`mcp`](https://pypi.org/project/mcp/) SDK
(`FastMCP`), plain JSON for storage.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues