MCP Recruiting Agent
It is an MCP recruiting server that lets you search jobs and candidates, retrieve details, score matches, and access hiring policies.
search_jobs: Search open jobs by keywords (skills/title) and optional location.
get_job: Fetch a single job posting by ID (e.g., J-101).
search_candidates: Rank candidates by skill overlap, with optional limit and minimum years.
get_candidate: Fetch a candidate profile by ID (e.g., C-201), excluding contact details.
score_match: Get a transparent skill/experience match score between a candidate and a job.
get_policy: Retrieve hiring policy text by topic (eeo, screening, data_privacy, ai_use).
Also exposes a
policy://{topic}resource and ashortlistprompt for MCP clients.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Recruiting AgentFind candidates for J-101 and score their matches"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Recruiting Agent
A guarded, multi-turn recruiting assistant: an MCP server, LangChain MCP adapters, a LangGraph agent, and conversational evals that include a guardrail ablation.
Evidence site: https://ingyukoh.github.io/mcp-recruiting-agent/ · Skill → code map: MATCHER_EVIDENCE.md
Measured result
13 multi-turn conversations (30 turns), run against the real MCP server over stdio. The same suite is also run with the guardrails removed.
Metric | Guardrails on | Guardrails off |
Conversation success | 100% | 61.5% |
Turn pass rate | 100% | 73.3% |
Tool-trajectory accuracy | 100% | 96.7% |
Block recall (injection + EEO requests) | 100% | 0% |
Turns leaking PII | 0 | 1 |
Turns echoing injected instructions | 0 | 1 |
These numbers come from python -m recruiting_agent.evals (report, per-turn JSON). Run the suite locally to reproduce them; scripts/compare_results.py checks a fresh report against the checked-in results.
Scope warning: this is a portfolio system running on synthetic data, not a client deployment. By default the agent uses a deterministic LangChain chat model (OfflinePlannerModel), so tests and evaluations are reproducible without API keys. I wrote the scenarios alongside that model, so the guarded 100% is a regression baseline, not a claim about general LLM quality. The informative comparison is the ablation. Set RECRUITING_AGENT_LLM=anthropic to run the same graph and evals with Claude.
Related MCP server: mcp-greenhouse
Architecture
flowchart LR
U[User turn] --> IG[input_guard<br/>injection · EEO · PII]
IG -->|blocked| END[Answer]
IG --> A[agent<br/>LangChain chat model]
A -->|tool_calls| T[ToolNode<br/>LangChain MCP tools]
T <-->|stdio| M[(MCP server<br/>6 tools · resource · prompt)]
T --> TG[tool_output_guard<br/>indirect injection]
TG --> A
A --> OG[output_guard<br/>PII · grounded ids]
OG --> ENDMCP server (
mcp_server.py):search_jobs,get_job,search_candidates,get_candidate,score_match,get_policy; apolicy://{topic}resource; ashortlistprompt. Uses stdio by default, or--httpfor streamable HTTP. Contact fields never leave the server.LangChain integration (
graph.py):MultiServerMCPClientandload_mcp_toolsturn MCP tools into LangChainBaseTools. One MCP session is kept open for the process lifetime.LangGraph agent: a
StateGraphwith guard nodes, aToolNodeloop capped at 4 rounds, andInMemorySavercheckpoints perthread_id, so follow-ups like "the second candidate" or "score her for J-106" resolve.Guardrails (
guardrails.py):Direct prompt injection is refused.
Filtering by protected characteristics is refused, while benign near-misses are allowed (for example "over 10 years of experience" or "can we ask about age?").
PII is redacted on the way in and on the way out.
Instruction-like text inside tool results is neutralized (the JSON stays parseable).
Any C-/J- id in the answer that no tool returned is removed.
Conversational evals (
evals.py,conversations.json): each turn is checked for tool trajectory, required and forbidden content, block decision, and guard events. Leak metrics are computed from the answer text independently of the guard's own log.Serving: FastAPI
POST /chatandGET /health, a Docker image with a health check.
Reproduce
python -m venv .venv && source .venv/bin/activate
pip install -e '.[dev]'
pytest -q # 27 tests, including end-to-end over real MCP stdio
python -m recruiting_agent.evals # regenerates results/
python scripts/build_site.py # regenerates docs/index.html from results/
uvicorn recruiting_agent.api:app --port 8000
curl -s localhost:8000/chat -H 'content-type: application/json' \
-d '{"thread_id":"demo","message":"Find candidates for J-101"}'Run with Claude instead of the offline model:
pip install -e '.[anthropic]'
export ANTHROPIC_API_KEY=... RECRUITING_AGENT_LLM=anthropic
python -m recruiting_agent.evals --variants guarded --out results-claudeUse the MCP server from any MCP client (for example Claude Desktop):
{"mcpServers": {"recruiting": {"command": "python", "args": ["-m", "recruiting_agent.mcp_server"]}}}Known limitations
Guardrails are pattern-based and deterministic, which makes them auditable but not exhaustive. A production system would add a classifier layer and red-team data.
Match scoring is a transparent skills/years formula, not a learned ranker. Every score is labelled as decision support, and a human decides.
The eval results reported here come from the offline model. Results with Claude have not been checked in.
langchain-mcp-adapters0.3.x requires the MCP Python SDK 1.x, somcp<2is pinned.
Available Tools
6 toolsget_candidateA
Get one candidate profile by id, e.g. C-201. Contact details are excluded.
| Name | Required | Description | Default |
|---|---|---|---|
| candidate_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does disclose one meaningful behavioral trait: contact details are excluded from the response. However, it does not mention not-found behavior, permissions, or whether the operation is strictly read-only, though 'Get' strongly implies it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The core action is front-loaded, and the additional note about contact details is relevant and concise. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema, so the description does not need to be extensive. It explains what is returned (candidate profile) and one exclusion (contact details), but without an output schema it leaves the full response shape and error behavior unspecified. This is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds an example ID format ('C-201') which gives the agent a concrete sense of the expected value. However, it does not explain the format rules, constraints, or what happens if the ID is invalid, leaving the parameter semantics only partially clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get one candidate profile by id'. It also gives an example ID format ('C-201') and clarifies scope by noting contact details are excluded. This clearly distinguishes it from sibling tools like search_candidates, which would be used for finding candidates rather than fetching a single one by ID.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when you already have a candidate_id and need a single profile. However, it does not explicitly mention when not to use it or point to search_candidates as the alternative for finding candidates without an ID. The usage context is clear but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_jobC
Get one job posting by id, e.g. J-101.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states that the tool retrieves a job posting by id, but does not disclose whether it is read-only, what the return format is, or any error/edge cases. This is a significant gap for a tool with no structured safety hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no unnecessary words, and the key information (the action and resource) is front-loaded. It is appropriately concise and well-structured, earning full marks for this dimension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema and no annotations, the description should include the return type, read-only status, and any caveats. It lacks all of these, making it incomplete for an agent to confidently invoke without further assumptions. For a simple tool it is still under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no description for job_id, and the tool description provides an example 'J-101', which gives a format hint but does not explain the semantics of the id or how to obtain it. This adds minimal value beyond the parameter name, and given 0% schema coverage, the description does not adequately compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Get' and the resource 'job posting', and provides an example id format 'J-101'. This distinguishes it from sibling tools like search_jobs (which searches) and get_candidate (different entity). However, it does not explicitly name a sibling to avoid confusion, so it doesn't fully earn a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description does not mention the context in which get_job should be preferred over search_jobs or other tools, leaving the agent to infer that it's for fetching a specific job by id. There is no explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_policyA
Get hiring policy text. Topics: eeo, screening, data_privacy, ai_use.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must fully disclose behavioral context on its own. It communicates read-only intent through the verb 'Get', but it does not state whether topics are exact enum values, what happens for invalid topics, whether all topics are available, or any other behavioral constraints. The list of topics gives some context but leaves important behavior unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. It front-loads the core action and then enumerates the relevant topics, making it easy to scan and process.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple one-parameter shape and the absence of an output schema, the description sufficiently identifies the resource and valid topics. However, it omits confirmation that the topic listing is exhaustive or that values must match exactly, and it provides no detail about the response format or potential error conditions, leaving a few practical gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only 'topic' as a string with no description and no enum, so the description adds meaningful value by listing the allowed topics: eeo, screening, data_privacy, and ai_use. This compensates for the 0% schema coverage, though it could be stronger by explicitly stating that these are exact permitted values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Get hiring policy text') and enumerates the available topics, making the tool's function immediately clear. It is clearly distinct from the sibling tools, which all concern jobs or candidates rather than policy content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance about when to use this tool versus alternatives, and no mention of which topics should be selected under which circumstances. The description merely implies the tool is for retrieving policy text, leaving usage decisions entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
score_matchC
Transparent skill/experience match score between a candidate and a job.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| candidate_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosing behavior. It says the score is 'transparent' but does not explain what that means operationally, whether the tool performs a read-only computation, or what the score represents (range, direction, criteria). This is a meaningful gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler or redundancy. It front-loads the core operation and the object of the operation. It is arguably under-specified, but that is a completeness concern, not a structure concern.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool, the description provides the essential purpose and output concept. However, with no output schema and no annotations, an agent is left without the score semantics, return format, or behavioral caveats needed for fully confident invocation. It is minimally adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must clarify the parameters. It does establish that one parameter refers to a candidate and the other to a job, which maps directly to candidate_id and job_id. However, it gives no additional detail about formats, constraints, or how the two IDs relate beyond the obvious matching purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a clear resource pair (candidate and job) and the output (a match score), so an agent can tell this tool computes a matching metric rather than merely retrieving a record. It does not explicitly contrast with siblings like get_candidate or get_job, but the function is distinct enough from search/get operations. 'Transparent' adds a specific quality of the score, slightly strengthening clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance about when to use this tool versus the siblings. The intended use is only implied by the tool name and the two identifiers. No exclusions, prerequisites, or alternative routing are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_candidatesB
Rank candidates by overlap with the given skills. Never returns contact details.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| skills | Yes | ||
| min_years | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses one key behavioral trait: 'Never returns contact details,' which is significant and prevents misuse. However, it does not mention other behaviors such as how results are ordered, what happens with zero matches, pagination, or any error conditions. The disclosure is minimal but at least includes a critical limitation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence that front-loads the core purpose and a critical constraint. Every word contributes value; there is no redundancy or fluff. It is appropriately sized for a tool with a simple, focused function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has three parameters, no annotations, and no output schema, the description is incomplete. It does not explain the meaning of 'limit' or 'min_years,' nor does it describe the response format beyond the absence of contact details. An agent would lack sufficient information to call the tool correctly (e.g., how to set limits, what output to expect) and would need to infer or seek additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate for parameter meaning. It only implicitly references the 'skills' parameter ('given skills'), but does not explain 'limit' or 'min_years.' The description adds no value for these parameters, leaving an agent to infer their meaning from the schema alone. Since it does clarify the primary input but ignores others, it partially compensates but falls short.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: 'Rank candidates by overlap with the given skills.' It uses a specific verb ('rank') and resource ('candidates'), and adds a distinguishing constraint ('Never returns contact details'). This clearly differentiates it from siblings like search_jobs (jobs vs candidates) and get_candidate (retrieval vs ranking).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool versus alternatives. It does not mention when to prefer search_candidates over search_jobs, get_candidate, or score_match. The only hint is the purpose itself, which implies usage for skill-based candidate ranking, but there is no 'use this when...' or 'for other searches, use...' statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_jobsA
Search open jobs by keywords (skills or title words) and optional location.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| location | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It only states the search criteria and doesn't disclose whether it's read-only, any limitations, pagination, or result format. The description is too minimal to convey behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the verb and resource. Every word earns its place; there is no fluff or redundancy. The structure is optimal for quick parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the basic operation and the two parameters, but given there is no output schema and no annotations, it lacks essential context like return format, pagination, or any usage constraints. While adequate for a simple search, it leaves the agent without enough detail to fully anticipate the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must clarify parameters. It explains that 'query' holds keywords (skills or title words) and 'location' is optional, adding semantic meaning. However, it omits details like keyword combination logic or location format, so it only partially compensates for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Search'), a specific resource ('open jobs'), and the criteria (keywords and optional location). This distinguishes it from siblings like get_job (fetch a single job) and search_candidates (search people).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by describing what the tool does, but it does not explicitly state when to choose this over search_candidates or get_job. No exclusions, alternatives, or conditions are provided, leaving the agent to infer from context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
get_candidate - First observed
get_job - First observed
get_policy - First observed
score_match - First observed
search_candidates - First observed
search_jobs
TDQS
Scored across 6 tools
Each tool targets a distinct action: searching vs retrieving jobs/candidates, matching a candidate to a job, and fetching policy. There is no overlap in purpose; even search vs get are clearly separated by scope (list vs single record).
All tool names follow a consistent verb_noun pattern with snake_case (e.g., search_jobs, get_candidate, score_match). The verbs are action-oriented and the nouns are domain objects, making the set predictable and easy to navigate.
With 6 tools, the server is well-scoped for a recruiting assistant: two search tools, two retrieval tools, a matching tool, and a policy lookup. Each tool serves a clear purpose without redundancy or bloat.
The surface covers the core recruiting workflow of discovering jobs and candidates, examining details, scoring matches, and consulting hiring policies. Since this appears to be a read-only analysis agent, no update/create/delete operations are required; the set has no obvious dead ends.
Maintenance
Related MCP Connectors
Query professional profiles, search candidates, and get AI-powered summaries and job fit analysis.
AI screening interviews, resume matching, and evidence-linked scorecards for recruiting teams.
AI agent recruiting: talent pool match, reference checks, credit packs via Stripe MPP
Recruiting tools for candidate sourcing, enrichment, ATS workflows, campaigns, and outreach.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables LLMs to access candidate information including resume, LinkedIn, GitHub, and contact via email.10 npm81MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to search candidates, view job postings, and manage applications in Greenhouse ATS via natural language queries.168 npmMIT
- AlicenseAqualityBmaintenanceEnables managing Recruitee/Tellent recruiting pipelines from Claude, including reading roles and candidates, creating candidates, and writing evaluations and notes with previews and safe write operations.14MIT
- AlicenseAqualityCmaintenanceEnables AI agents to search, rank, and explain job matches through a secure read-only interface that blocks prompt injections and unsafe content.4MIT