Skip to main content
Glama
hanjot

Job Application Tracker

by hanjot
README.md
# Job Application Tracker (MCP Server)

![CI](https://github.com/hanjot/job-application-tracker-mcp/actions/workflows/ci.yml/badge.svg)
![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)

**[Live dashboard →](https://hanjot.github.io/job-application-tracker-mcp/)**

A real [Model Context Protocol](https://modelcontextprotocol.io) server that
turns my own job search — 200 real applications sent between August 10 and
September 29, 2026 — into something an AI assistant can query and update in
plain language.

This is a follow-up to my first MCP project
([mcp-filesystem-connection](https://github.com/hanjot/mcp-filesystem-connection)),
which proved the basic AI-to-local-file connection. This one is built on my
actual job-search data and does real read/write work: querying, filtering,
and updating applications through MCP tools, not just reading a file.

## Why I built it — and what it's actually for

I'm a Lead Technical Program Manager currently in an active job search, and
by the time I built this I had sent 150+ applications with no single place
to see them — just a folder of PDFs named inconsistently
(`Company_Role_Date.pdf`, `RoleCompanyDate.pdf`, some with location, some
without).

This tracker is deliberately **not** about callbacks or interview status.
The question I actually care about is: across everything I've applied to,
what skills is the market asking for — Salesforce, SAP, cybersecurity, AI,
and to what depth — and what is each of those actually paying? Seeing that
pattern across 200 applications, instead of one job description at a time,
is the actual value.

## What it does

`parse_applications.py` reads the raw filenames and file timestamps from my
Applications folder and parses each one into: company, role, applied date,
a `skills_required` tag list (Salesforce, SAP, AI/ML, AI/GenAI,
Cybersecurity, Cloud, Data/Analytics, Compliance, Payments, Agile, Data
Privacy, and so on), and — where a pay range is either stated in a saved
posting or known from a real conversation about that specific posting — a
`salary_range`.

`server.py` is the MCP server. It loads `applications.json` (the parsed
output) and exposes these tools to any MCP client:

| Tool | What it does |
|---|---|
| `list_applications` | Filter by company, by required skill, or by date |
| `get_application` | Full detail on one application by id |
| `add_application` | Log a new application, with the real job link + job description text if available (auto-tags skills and extracts a salary range from the text when present) |
| `update_application` | Correct company/role/notes, add a real `job_url`/`job_description` retroactively (auto re-tags `skills_required` and re-extracts `salary_range` from the real text), or set skills/salary by hand |
| `get_summary` | Totals, date range, repeat companies, and — the main point — **skills_frequency**: how often each skill/technology showed up across every posting |
| `get_salary_summary` | Average disclosed pay by skill category (e.g. AI vs. Cybersecurity vs. Salesforce) — hourly rates annualized so everything compares on the same basis |
| `get_course_priority` | Ranks which course/certification to prioritize next, based on real demand across all 200 applications (pulls from my existing AI Course Priority Plan where a matching course exists) |
| `find_duplicates` | Companies applied to more than once, with each application listed |

`skills_report.html` is a standalone chart + table view of the same data
(open it in any browser) — a horizontal bar chart of skill frequency across
all 200 applications, a second chart of average disclosed salary by skill
category, and the course-priority ranking below both.

**Real job descriptions, going forward:** each application also has a
`job_url` and `job_description` field. When a real posting's text is saved
(via `add_application` or `update_application`), `skills_required` is
automatically re-tagged from that actual text instead of guessed from the
title, and `salary_range` is extracted automatically if the text states a
pay range — the same keyword matcher, just run against real content. This
is how new applications get added from here on: job link + full JD text in,
accurate skill tags and pay data out.

## Real numbers from my own search (as of Sept 29, 2026)

- **200 applications** tracked, spanning **Aug 10 – Sept 29, 2026**, across
  **133 unique companies**
- Skills flagged from job titles (and real posting text where saved) so
  far: **Program/Project Management** (135), **AI/GenAI/ML** (22),
  **Cybersecurity** (11), **Data/Analytics** (9),
  **Compliance/Risk/Governance** (6), **Infrastructure/DevOps** (5), plus
  smaller counts for Agile/Scrum, Cloud, Payments, Salesforce, and Data
  Privacy
- 48 applications have no skill tag yet — their titles didn't contain a
  recognizable keyword (see limitation below)
- **35 of 200 applications have a disclosed salary range** — 2 from a real
  saved job description, the rest from pay ranges discussed for that
  specific posting. By skill category, average annualized base pay ranges
  from roughly **$190K (AI/GenAI/ML)** to **$243K (Cybersecurity)** among
  the categories with disclosed data; see `get_salary_summary` or the
  dashboard for the full breakdown. Categories with zero disclosed data
  simply don't appear — nothing here is a guessed or averaged-out figure.

## Running it

```bash
pip install -r requirements.txt
python parse_applications.py      # rebuilds applications.json from raw_listing.json
mcp dev server.py                 # opens the MCP Inspector to try the tools by hand
```

To connect it to an MCP-compatible client, add it as a stdio server that
runs `python server.py` from this folder.

## Verifying it's real

Two layers of testing, both run automatically on every push via GitHub
Actions (see the CI badge above):

1. **Unit tests** (`tests/test_parse_applications.py`, run with `pytest`) —
   check the filename-parsing logic against real, tricky cases from the
   actual dataset: camelCase titles with no separators, a company name
   ("Marketing") that contains a month abbreviation as a substring
   ("mar"), dates that must fall back to the file's save time when the
   filename has none, self-authored resume filenames that start with my
   own name, flagged duplicate files, and salary-range text extraction.
2. **End-to-end protocol test** (`test_client.py`) — spawns `server.py` as
   an actual MCP server over stdio, calls `initialize`, lists the tools,
   and calls `get_summary`, `get_salary_summary`, `list_applications`,
   `add_application`, `update_application`, `get_course_priority`, and
   `find_duplicates` — reading back real results and confirming the write
   tools actually persisted changes to `applications.json`. That's a
   genuine protocol round trip, not a mocked call.

Run both locally with:

```bash
pip install -r requirements.txt pytest
python parse_applications.py
python -m pytest tests/ -v
python test_client.py
```

## Honest limitations

- **For most of the 200 applications parsed from filenames,
  `skills_required` is still inferred from the job title only** — I don't
  have the original posting text saved for those, so it's a best-effort
  keyword match against the role text, not a transcription of each
  posting's actual requirements. Going forward, any application added with
  a real `job_description` gets accurate, text-based tags instead (see
  above); I'm backfilling the older ones with `update_application` as I
  revisit real postings.
- **Salary data is partial and curated, not exhaustive.** Only 35 of 200
  applications have a disclosed range — either extracted from a saved job
  description, or entered by hand from a specific pay conversation I had
  about that posting. I deliberately did not guess or average a figure for
  postings without a real disclosed number, so most applications (and a
  few whole skill categories) simply have no salary data yet.
- The filename parser is heuristic in general. A small number of
  applications (roughly 5%) come through with an unclear role or company
  because the original filename ran words together with no separator.
- There is no status/callback field by design — this tracker is about
  required skills (and pay) across the whole search, not per-application
  outcomes.
- Four files that were saved copies of my base resume, or duplicate
  `.docx` copies of a `.pdf` I'd already counted, are excluded from the
  counts rather than counted as applications.

## Stack

Python, the official [`mcp`](https://pypi.org/project/mcp/) SDK
(`FastMCP`), plain JSON for storage.