Job Search MCP
职位搜索引擎
在5家大型制药公司中快速、可靠地搜索职位。
直接从Amgen、Bayer、GSK、Novartis和Pfizer的职业网站搜索职位。即时提取职位标题、描述、要求和申请链接。
演示: https://[your-vercel-app].vercel.app/
功能特性
✅ 同时搜索5家公司
✅ 提取详细的职位信息(标题、描述、要求、截止日期)
✅ 使用错误代码跟踪失败的提取(404、超时等)
✅ 生成包含结果的CSV报告
✅ 零外部依赖,基于正则表达式的快速解析
✅ TypeScript + 严格的类型安全
✅ 全面的错误处理和分类
✅ 限速API(每IP每天5次搜索)
Related MCP server: trackly-cli
技术栈
前端: React + TypeScript + Tailwind CSS
后端: Next.js + Node.js
解析: 基于正则表达式的HTML提取(无重型依赖)
运行时: Node.js(v18+)
语言: TypeScript 5.3+
构建: TypeScript编译器(tsc)
部署: Vercel(推荐)或AWS Lambda
快速开始
尝试演示
访问:https://[your-vercel-app].vercel.app/
您将被重定向到职位搜索界面。输入职位标题,选择公司,并即时浏览结果。
本地开发
# Install dependencies
npm install
# Build TypeScript
npm run build
# Run a test
node dist/test/test-manager-jobs.js
# Start development server (requires Next.js setup)
npm run dev生产部署
# Deploy to Vercel (recommended)
npm install -g vercel
vercel
# Or deploy to AWS
# See docs/DEPLOYMENT.md for AWS Lambda setup项目架构
User searches for jobs → Demo page (/app/demo/page.tsx)
↓
React UI Component
- Search input
- Company multi-select
- Sortable results tables
↓
REST API (/api/search-jobs)
↓
┌───────────────┬────────────────┬────────────────┐
│ │ │ │
Amgen Bayer GSK Novartis Pfizer
(Workday) (Eightfold AI) (Workday) (Drupal) (Workday)
│ │ │ │
└───────────────┴────────────────┴────────────────┘
↓
Search Executor (src/search-executor.ts)
- Fetches job URLs from each site
- Parses HTML for job listings
↓
Extractor Registry (src/extractors/)
- Extracts job details from each URL
- Company-specific parsers
- Error tracking & classification
↓
Extraction Helpers (src/extraction-helpers.ts)
- CSV report generation
- Error aggregation
↓
REST API Response (JSON)
↓
Demo Page displays results
- Success table: Jobs with details
- Error table: Failed extractions
- Download CSV buttons配置
src/config.json 是项目支持的公司信息的唯一来源。
示例:
{
"projectname": "Job Search MCP",
"sites": [
{
"name": "Amgen",
"search_url": "https://amgen.wd1.myworkdayjobs.com/Careers?q=Engineer"
},
{
"name": "Pfizer",
"search_url": "..."
}
]
}添加新公司时,只需提供公司名称和一个可用的搜索URL。
站点定义
每家公司由 sites/ 下的一个单独文件表示。
例如:
sites/amgen.json结构必须遵循 test/sample.json。
站点定义包含:
公司名称
职业网站URL
搜索URL
支持的搜索参数
参数标签
参数类型
可用的参数值
参数结构有意设计为数组而不是固定的JSON键,因为不同的职业网站暴露不同的搜索参数。
例如,一个站点可能暴露:
location
country
jobType而另一个站点可能暴露:
location
timeType
LocationCountry
jobFamilyGroup
workerSubTypeMCP不得假设每家公司都支持相同的参数。
职位提取器
项目包含特定于站点的职位提取器,用于解析单个职位发布URL并提取详细信息。
提取的数据
每个提取器获取:
职位标题 - 职位名称
职位描述 - 完整的职位描述/职责(不包括页眉/页脚)
资格要求 - 要求、资质和技能
截止日期 - 申请截止日期(YYYY-MM-DD格式,如无则为空)
申请链接 - 直接申请URL(可能与职位发布URL不同)
可用的提取器
src/extractors/
├── types.ts # JobExtractor interface & types
├── amgen.ts # Amgen (Workday-based)
├── pfizer.ts # Pfizer (Workday-based)
├── bayer.ts # Bayer (Eightfold AI)
├── gsk.ts # GSK (Workday-based)
├── novartis.ts # Novartis (Drupal)
└── index.ts # ExtractorRegistry使用示例
import { ExtractorRegistry } from './src/extractors/index.js';
const registry = new ExtractorRegistry();
const amgenExtractor = registry.getExtractor('amgen');
const result = await amgenExtractor?.extract(
'https://amgen.wd1.myworkdayjobs.com/job/India---Hyderabad/Assoc-Director---Data-Product-Mgmt_R-219150'
);
if (result?.success && result.data) {
console.log(result.data.jobTitle);
console.log(result.data.jobDescription);
console.log(result.data.eligibility);
}测试
项目包含一套全面的测试套件,用于验证搜索和提取功能。
测试套件概览
所有测试都是独立的TypeScript文件,可以独立运行:
npm run build
node dist/test/[test-name].js可用的测试
1. test-config.ts - 配置加载测试
测试公司配置是否正确地从 src/config.json 加载。
node dist/test/test-config.js目的: 验证配置结构和公司发现 输出: 列出可用的公司及其搜索URL
2. test-search.ts - 职位搜索测试
测试所有公司的搜索功能。
node dist/test/test-search.js目的: 验证搜索是否返回有效的职位URL 输出: 每家公司"Manager"职位的搜索结果 注意: 需要互联网连接以访问实际的职业网站
3. test-extractors.ts - 职位提取测试
测试每家公司职位URL的职位详情提取功能。
node dist/test/test-extractors.js目的: 验证职位标题、描述和资格要求的提取 输出: 提取成功率和字段详情 注意: 需要来自test-search.ts输出的真实职位URL
4. test-manager-jobs.ts - 端到端集成测试
完整的流水线测试:搜索职位 → 提取详情 → 生成报告
node dist/test/test-manager-jobs.js目的: 完整的集成测试,包含错误跟踪和CSV报告生成 输出:
test/manager-jobs-success.csv- 成功提取的职位数据test/manager-jobs-errors.csv- 提取错误(404、超时等)显示成功率和错误分类的控制台摘要
运行所有测试
npm run build
node dist/test/test-config.js
node dist/test/test-search.js
node dist/test/test-extractors.js
node dist/test/test-manager-jobs.js测试输出文件
生成的CSV报告存储在 test/ 文件夹中:
manager-jobs-success.csv- 成功的职位提取manager-jobs-errors.csv- 带有错误代码的失败提取尝试用于调试的示例HTML文件
这些文件在测试运行期间生成,可以安全删除。它们位于 .gitignore 中。
test/sample.json
test/sample.json 定义了单个公司文件的预期结构。
它是一个按示例/模板定义的schema,而不是公司注册表。
当前示例使用诸如 location、timeType、LocationCountry、jobFamilyGroup 和 workerSubType 等参数。
BUILD.md
BUILD.md 包含用于从公司配置构建MCP的AI/开发流程说明。
构建过程应:
读取
src/config.json。处理
sites中列出的每家公司。访问/分析提供的搜索URL。
确定公司实际的职业/搜索结构。
发现可用的搜索参数及其值。
生成或更新相应的
sites/<company>.json。确保生成的文件遵循
test/sample.json定义的结构。构建/更新通用的MCP实现。
验证所有配置的站点都可以被搜索。
UPDATE.md
参见 ai/UPDATE.md 了解在创建新版本时重建项目的说明。
当 src/config.json 发生变化时,AI必须重建所有公司定义,而不仅仅是新添加的公司。
这是有意为之。
现有的职业网站可能会更改其:
搜索URL
查询参数
筛选器名称
筛选器值
职业网站结构
ATS实现
因此,每个版本都应重新检查现有的 sites/*.json 文件与当前在线职业网站的一致性。
src/config.json updated
│
▼
Rebuild ALL sites
│
├── New company → create site JSON
│
└── Existing company → re-analyze and update
│
▼
Rebuild common MCP
│
▼
Validate
项目结构
JobSearchMCP/
├── src/ # Source code & configs
│ ├── server.ts # MCP server entry point
│ ├── search-executor.ts # Search execution & parsing
│ ├── config-loader.ts # Configuration loader
│ ├── types.ts # TypeScript types
│ ├── config.json # Company registry
│ ├── site_configurations.json
│ └── site_analysis.json
├── sites/ # Company-specific configs
│ ├── amgen.json
│ ├── pfizer.json
│ ├── novartis.json
│ ├── bayer.json
│ └── gsk.json
├── test/ # Tests & test data
│ ├── test-*.js # Test scripts
│ ├── sample.json # Configuration template
│ └── *.html # Sample HTML files
├── ai/ # AI development notes (Gitignored)
│ ├── AI.md
│ └── UPDATE.md
├── reports/ # Documentation
│ ├── IMPLEMENTATION.md
│ ├── ANALYSIS_GUIDE.md
│ ├── MCP_USAGE.md
│ └── MIGRATION.md
├── dist/ # Compiled JavaScript
├── package.json # Dependencies & scripts
├── tsconfig.json # TypeScript config
└── README.md # This file技术
运行时:Node.js 语言:TypeScript MCP SDK:官方Model Context Protocol TypeScript SDK 配置:JSON
设计原则
项目将特定于站点的知识与通用MCP逻辑分离。
sites/*.json
= How a particular company career site works
MCP implementation
= How to search any configured company
AI
= Understand the user's request and select/use the appropriate
company search configurationMCP不应包含关于 location、remote、full_time 或 job_type 等参数的硬编码假设。
只有当该公司的职业网站实际支持该参数或暴露配置所需的信息时,该参数才对该公司存在。
目标
目标是创建一个可复用的职位搜索MCP,其中添加公司主要是将它们的搜索URL添加到 config.json 的问题,让AI构建过程自动发现和维护特定于站点的配置。
This server cannot be deployed
Maintenance
Related MCP Connectors
Public MCP server for discovering open jobs. Search, filter, and get application links.
GetJobzi MCP server for job search, application tracking, and career forecasting.
JobsPipe — data pipeline of every job posting on the web. Search live, normalized job postings from 30+ ATS feeds and job boards for AI agents via MCP.
AI job search MCP — fact-checked jobs, application tracker, alerts. ChatGPT, Claude, Cursor.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceMCP server that exposes job search data from multiple boards, enabling clients to query and manage job listings via natural language.7MIT

trackly-cliofficial
AlicenseNot gradedqualityAmaintenanceMCP server for job search and application tracking, enabling AI agents to search jobs, get details, manage applications, and find contacts across 128K+ jobs and 1,900+ companies.258 npm3MIT- AlicenseCqualityCmaintenanceAn MCP server that enables AI-assisted job search workflows including job discovery, application tracking, resume evaluation, and cover letter generation, with support for multiple job sources and scheduled scraping.8342 npm1AGPL 3.0
- AlicenseNot gradedqualityDmaintenanceA custom MCP server that exposes a jobs database to any MCP-compatible LLM client, allowing users to ask in plain English to search, filter, and match job openings.MIT