careerproof-mcp
Provides tools for indexing GitHub repositories and extracting verified evidence such as commits, pull requests, architecture docs, and dependency manifests for interview preparation and job-requirement matching.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@careerproof-mcpAnalyse this Solutions Architect JD and map my GitHub projects to each requirement."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
careerproof-mcp
Evidence-backed interview preparation through the Model Context Protocol.
CareerProof turns a candidate's real project history — GitHub repositories, a CV, and project notes — into interview preparation that is traceable to specific evidence: commits, pull requests, architecture docs, and dependency manifests. Every answer it helps produce distinguishes between:
Verified evidence — pulled directly from GitHub
Reasonable inference — evidence added manually, not independently verified
User-provided claim — extracted from the candidate's own CV
Missing evidence — explicitly flagged, never silently invented
Why
Ask a connected MCP client:
Analyse this Solutions Architect job description and show me where my GitHub projects prove each requirement.
and get back a requirement-by-requirement match, each one backed by a cited source:
Job requirement | Match | Supporting evidence |
Power Platform | Strong | CertMate workflow documentation |
API architecture | Strong | REST integration and service-layer code |
SQL | Strong | Database schema and stored procedures |
CI/CD | Weak | GitHub Actions exists but deployment evidence is limited |
Team leadership | Unproven | No evidence found in the supplied sources |
See docs/demonstration.md for a full walkthrough and
examples/sample-evidence-report.json
for a complete sample response.
Related MCP server: github-repo-intel-mcp
Design principle: no LLM in v1
CareerProof deliberately does not call an LLM API. Job description parsing
and requirement matching use bullet parsing, a curated keyword dictionary, and
token-overlap scoring. STAR-answer generation builds a cited outline, not
prose. The connected MCP client (Claude Desktop, VS Code Copilot, etc.)
supplies the language reasoning on top of this server's structured, traceable
evidence — which keeps the server cheap, private, and easy to self-host.
See docs/architecture.md for the full rationale.
MCP primitives
Tools (11) — actions such as indexing repositories, analysing job descriptions, matching requirements, generating STAR outlines, and exporting a preparation pack:
careerproof_add_candidate_profile, careerproof_index_repository,
careerproof_add_project_evidence, careerproof_analyse_job_description,
careerproof_match_requirements, careerproof_find_evidence,
careerproof_generate_star_answer, careerproof_find_evidence_gaps,
careerproof_generate_interview_questions, careerproof_score_interview_answer,
careerproof_export_preparation_pack
Resources (5) — read-only, application-controlled views:
careerproof://candidate/profile, careerproof://jobs/{jobId},
careerproof://projects/{projectId}, careerproof://evidence/{evidenceId},
careerproof://skills/matrix
Prompts (5) — reusable, host-surfaced workflows:
prepare_for_interview, create_star_answer, challenge_cv_claim,
run_mock_technical_interview, identify_portfolio_gaps
Quick start
npm install
npm run build
npm start # runs dist/server.js over stdioOr run directly from source during development:
npm run devTry it with the MCP Inspector:
npx @modelcontextprotocol/inspector npx tsx src/server.tsRegister with an MCP host
Point your host at the built server (see mcp.json for a
ready-made config):
{
"mcpServers": {
"careerproof": {
"command": "node",
"args": ["dist/server.js"],
"env": { "CAREERPROOF_DB_PATH": "./data/careerproof.db" }
}
}
}Environment variables
Variable | Purpose | Default |
| Path to the local SQLite database |
|
| Optional GitHub token for higher API rate limits / private repos | unset (public, unauthenticated) |
Example tool call
{
"tool": "careerproof_generate_star_answer",
"arguments": {
"competency": "Describe a time you designed a complex solution",
"project": "CertMate EICR",
"maximumWords": 250
}
}{
"answer": {
"situation": "...",
"task": "...",
"action": "...",
"result": "..."
},
"confidence": 0.84,
"evidence": [
{ "source": "docs/architecture.md", "lines": "18-42", "type": "verified" }
],
"missingInformation": [
"No measurable performance improvement was documented"
]
}Repository structure
careerproof-mcp/
├── src/
│ ├── server.ts # entry point (stdio transport)
│ ├── tools/ # 11 MCP tools
│ ├── resources/ # 5 MCP resources
│ ├── prompts/ # 5 MCP prompts
│ ├── github/ # GitHub REST client + repository indexer
│ ├── evidence/ # CV parsing, evidence store, export pack
│ ├── matching/ # requirement extraction, matching, STAR/gap/question logic
│ └── database/ # Drizzle schema + SQLite client
├── examples/ # sample CV, job description, evidence report
├── evals/ # judgement evals for the matching logic
├── tests/ # Vitest unit + MCP protocol tests
├── docs/ # architecture, threat model, demonstration walkthrough
├── Dockerfile
├── mcp.json
└── README.mdDevelopment
npm test # vitest unit + protocol tests
npm run lint # tsc --noEmit
npx tsx evals/requirement-matching.eval.ts # judgement evalDocker
docker build -t careerproof-mcp .
docker run -i -v careerproof-data:/data careerproof-mcpThe server speaks stdio, so docker run -i (interactive, no TTY) is how an
MCP host would launch it as a subprocess.
Roadmap
Streamable HTTP transport + OAuth for a remotely-hosted deployment
Optional local embeddings for semantic evidence search
PDF CV import
Docs
License
MIT
Available Tools
11 toolscareerproof_add_candidate_profileAdd candidate profileA
Registers a candidate's CV (Markdown or plain text) and extracts individually-citable claims from it. Every extracted claim is stored with confidence 'user_claim' — it is what the candidate says about themselves, not independently verified.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Candidate's full name | |
| cvText | Yes | Full CV content in Markdown or plain text | |
| summary | No | Optional short profile summary | |
| headline | No | Short professional headline, e.g. 'Solutions Architect' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavioral context beyond annotations: extracted claims are tagged with confidence 'user_claim' and are unverified assertions by the candidate, which tells the agent how to treat downstream results. Annotations only cover readOnly/idempotent/destructive hints; the semantic 'not verified' caveat is valuable. Doesn't disclose return shape or duplicate-profile behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the primary action and followed by the critical semantic caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotation set (readOnly=false, idempotent=false, destructive=false) matches a registration tool, and the description covers the write action plus claim semantics. No output schema exists, yet the description explains the key return concept (individually-citable claims with confidence). Lacks explicit 'what happens on duplicate candidate' detail, hence not 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are fully documented in the schema (name, cvText, summary, headline). The description adds only format guidance ('Markdown or plain text'), which the schema already states, so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Registers') and resource ('a candidate's CV') plus the generative action ('extracts individually-citable claims'). Far clearer than siblings like careerproof_add_project_evidence, which the description implicitly distinguishes by focusing on the candidate's own CV.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage (onboarding a candidate's CV before scoring/matching), but never explicitly says when to use this vs careerproof_add_project_evidence or careerproof_index_repository, nor any prerequisites. Adequate but no alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_add_project_evidenceAdd manual project evidenceA
Manually attaches an evidence item to an indexed project (e.g. an architecture decision explained in conversation). Defaults to 'inferred' confidence since it wasn't pulled directly from a verifiable source.
| Name | Required | Description | Default |
|---|---|---|---|
| excerpt | Yes | The evidence text itself | |
| lineEnd | No | ||
| filePath | No | ||
| lineStart | No | ||
| projectId | Yes | Project ID returned by careerproof_index_repository | |
| confidence | No | Defaults to 'inferred' | |
| sourceType | No | manual |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the write/safety profile is covered structurally. The description adds the useful rationale for the default confidence value, but says nothing about the consequences of non-idempotency (e.g., whether repeated calls create duplicate evidence rows) or what the caller gets back, which is the kind of extra context a mutation tool benefits from.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loaded with the verb and resource before the rationale, with no wasted framing. The second sentence slightly duplicates the schema's declared default, but it earns its place by explaining why the default exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description is the only place an agent could learn what the call returns, and it is silent on that. It also leaves the spatial/line parameters and sourceType semantics unaddressed, though for a straightforward add operation whose safety profile is covered by annotations, it is minimally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43% across 7 parameters, so the description should compensate, but it does not: filePath, lineStart, lineEnd, and sourceType are unexplained in both places. The only parameter reference is the 'inferred' confidence default, which merely restates the schema's own description and enum default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (manually attaches), the resource (an evidence item), and the target (an indexed project), with a concrete example clarifying what counts as manual evidence. The word 'manually' plus the 'indexed project' reference cleanly separates it from automatic extraction siblings like careerproof_index_repository and from retrieval siblings like careerproof_find_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mention of a default 'inferred' confidence because the item 'wasn't pulled directly from a verifiable source' implies this tool is for evidence that cannot be sourced automatically, which hints at when to use it. However, it never explicitly says when to prefer this over careerproof_index_repository or other add_* siblings, nor any prerequisites beyond the implicit need for an indexed project.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_analyse_job_descriptionAnalyse a job descriptionA
Stores a pasted job description and extracts structured requirements (skills, tools, competencies) using bullet parsing and a curated keyword dictionary. No LLM is used here — the connected client can add further reasoning on top of these structured requirements.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Job title, e.g. 'Solutions Architect' | |
| descriptionText | Yes | Full job description text |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it discloses the extraction mechanism (bullet parsing plus a curated keyword dictionary), that no LLM is involved (deterministic, limited recall), and that the JD is persisted. It still omits what happens on re-submission, whether the text is deduplicated, and any permission or size constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and followed by the mechanism and the important 'no LLM' caveat. No filler, though the em-dash clause could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no output schema, the description adequately covers intent, mechanism, return shape (structured skills/tools/competencies), and the determinism caveat. What is missing — error behavior, persistence semantics, and downstream pipeline ordering — is minor but non-trivial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented there, so the baseline is 3. The description adds no syntax, format, or length guidance beyond the schema (e.g., it does not mention the 20-character minimum).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair ('stores a pasted job description and extracts structured requirements') and even names the extracted categories (skills, tools, competencies). It does not explicitly differentiate itself from siblings like careerproof_match_requirements, but an agent can tell it is the ingestion/extraction step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: storing and extracting a JD is clearly the upstream step before match_requirements or find_evidence_gaps, and the note that 'the connected client can add further reasoning' hints at pipeline position. No explicit when-to-use, when-not-to-use, or named alternative is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_export_preparation_packExport interview preparation packB
Assembles the complete evidence-backed preparation pack for a job: requirement matches, evidence gaps, and generated interview questions. Returns either JSON or a Markdown report.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | Job ID returned by careerproof_analyse_job_description | |
| format | No | json | |
| candidateId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose the output medium ('Returns either JSON or a Markdown report'), which is genuinely useful. However, it says nothing about side effects (is anything persisted?), cost/latency implied by generating interview questions, or whether it requires prior analysis state, so the disclosure is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences: the first front-loads what is assembled and its components, the second covers the return format. No filler, no restatement of the tool name, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter aggregation tool with no output schema and no annotations, the description covers purpose and output medium but leaves the candidateId parameter unexplained and gives no usage routing against the many sibling tools that perform the same sub-tasks. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description should compensate, and it partially does by explaining the output format choice (JSON vs Markdown) that the 'format' parameter controls. It adds nothing about 'candidateId' (unlabeled in schema and absent from the description) and only implicitly covers jobId via the task framing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Assembles') and a well-defined artifact ('complete evidence-backed preparation pack for a job'), then enumerates the three contents: requirement matches, evidence gaps, and generated interview questions. That enumerates capabilities that map onto sibling tools (match_requirements, find_evidence_gaps, generate_interview_questions), implicitly distinguishing this aggregator from them, but it never names those siblings or states the distinction explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance: nothing says whether to reach for this aggregate pack versus calling careerproof_match_requirements, careerproof_find_evidence_gaps, and careerproof_generate_interview_questions individually, nor whether a candidate profile must exist first. The only prerequisite information (jobId coming from careerproof_analyse_job_description) lives in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_find_evidenceFind evidenceA
Free-text search across all stored evidence (commits, pull requests, docs, dependency manifests, CV claims and manual notes). Every result carries its confidence tag (verified / inferred / user_claim / missing) so the caller knows how much to trust it.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | Text to search for in evidence excerpts and file paths | |
| projectId | No | Restrict the search to one project |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose a valuable behavioral trait not in the schema: every result carries a confidence tag (verified/inferred/user_claim/missing), which tells the caller how much to trust hits. However, it says nothing about ranking, pagination, result volume, or performance characteristics of a cross-corpus search.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the scoping statement (all stored evidence) is front-loaded ahead of the return-value note. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description usefully covers the return shape (results tagged with confidence levels), which is the main thing an agent needs to interpret output. It is slightly incomplete on the undocumented limit parameter and on how results are ordered when many match, but is otherwise solid for a 3-param search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% – query and projectId are documented in the schema, while limit is not. The description adds no parameter-level detail beyond the schema (no ranking, ordering, or limit defaults), so it neither compensates for the undocumented limit nor enriches the documented ones. Baseline 3 for mid-range coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Free-text search across all stored evidence') and enumerates the searched corpora (commits, PRs, docs, manifests, CV claims, notes), which is far more than a restatement of the name. It does not explicitly distinguish itself from the sibling careerproof_find_evidence_gaps, so an agent must infer the boundary rather than read it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a discovery use case (searching stored evidence) but never states when to reach for this tool versus find_evidence_gaps, match_requirements, or the add_* tools, and gives no exclusions or prerequisites. Usage is inferable from the verb but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_find_evidence_gapsFind evidence gapsA
Re-runs requirement matching for a job and returns only the requirements with weak or unproven support, each with a concrete suggestion for what evidence to add.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | Job ID returned by careerproof_analyse_job_description | |
| projectId | No | ||
| candidateId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that matching is re-run (implying recomputation/refresh of prior results) and describes the shape of the return (only failing requirements, each with a suggestion). However, it omits side effects, whether prior results are overwritten, required permissions, or whether a prior analyse_job_description/match_requirements run is mandatory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the core action and the distinguishing output scope front-loaded. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain return values; it does so partially (gaps plus a concrete suggestion for each). But with no annotations and two of three parameters undocumented, the definition stops short of fully equipping an agent to call the tool correctly in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% — jobId is documented in the schema (and echoed indirectly by 'for a job'), while projectId and candidateId are entirely undocumented in both schema and description. The description adds no meaning beyond the schema and does not compensate for the two unexplained optional parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('re-runs requirement matching for a job') and precisely narrows the output scope ('only the requirements with weak or unproven support'), which distinguishes it from the sibling careerproof_match_requirements and careerproof_find_evidence without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the phrase 're-runs requirement matching' suggests this follows a prior analysis and complements match_requirements, but the description never explicitly says when to choose this tool over match_requirements or find_evidence. No prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_generate_interview_questionsGenerate interview questionsB
Generates likely interview questions from a job's requirements, prioritising requirements with the weakest evidence so preparation time is spent where it matters most.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| jobId | Yes | Job ID returned by careerproof_analyse_job_description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a real behavioral trait: questions are biased toward requirements with the weakest evidence. It says nothing about determinism, output format, limits on generation, or whether count is capped at the schema maximum of 30.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, leading with the action and following with the prioritisation rationale. Slightly dense and lacks any structural separation of purpose vs. behaviour, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter generation tool with no output schema, the description covers what it produces but omits the relationship to prerequisite tools, the meaning of `count`, and any indication of return shape or determinism. Adequate for selection, thin for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: jobId has a description pointing at careerproof_analyse_job_description, while `count` (1-30) is undocumented in both schema and description. The description never references either parameter, so it does not compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (generates) and resource (interview questions) and adds scope ('from a job's requirements'). It is clearly distinguishable in substance from siblings like careerproof_generate_star_answer or careerproof_score_interview_answer, though it never names or contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no named alternative. The phrase 'prioritising requirements with the weakest evidence' hints at intent but does not state prerequisites (presumably a prior careerproof_analyse_job_description) or when another sibling should be chosen instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_generate_star_answerGenerate a cited STAR-answer outlineA
Builds a Situation/Task/Action/Result outline for a competency and project, scaffolded entirely from stored evidence (commits, PRs, docs). Never invents narrative detail: gaps are listed explicitly in 'missing_information' rather than filled in. The connected client should turn this outline into flowing prose.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | The indexed project name to draw evidence from | |
| competency | Yes | The competency or question to answer, e.g. 'Describe a time you designed a complex solution' | |
| maximumWords | No | Target maximum length for the final prose answer |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a meaningful behavioral trait: it 'Never invents narrative detail' and surfaces gaps via 'missing_information.' That is valuable anti-hallucination context. It omits prerequisites (e.g., project must be indexed) and any auth/rate-limit context, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four compact sentences, front-loaded with what the tool builds, followed by its non-invention guarantee and the downstream handoff. No sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a generation tool with no output schema, the description usefully names a return field ('missing_information') and the outline structure, which compensates for the absent output schema. It stops short of stating prerequisite state (project indexed, evidence added) that an agent would need to call it successfully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters with clear semantics. The description reinforces the competency/project pairing but adds no syntax or formatting detail beyond the schema, and maximumWords is not addressed in prose. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Builds) and resource (Situation/Task/Action/Result outline), names its inputs (competency, project), and specifies the evidence source (commits, PRs, docs). This clearly distinguishes it from siblings like find_evidence or generate_interview_questions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes the pipeline role explicitly: 'The connected client should turn this outline into flowing prose,' which tells the agent where this tool sits relative to downstream work. However, it does not contrast with siblings like score_interview_answer or find_evidence, so the when-not-to-use case is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_index_repositoryIndex a GitHub repositoryB
Indexes a public (or token-accessible) GitHub repository: recent commits, pull requests, README/architecture docs, and dependency manifests. Everything pulled directly from GitHub is stored as 'verified' evidence because it is directly observed, not inferred.
| Name | Required | Description | Default |
|---|---|---|---|
| repoUrl | Yes | GitHub repository URL or 'owner/repo' shorthand | |
| candidateId | No | Candidate this project belongs to | |
| commitLimit | No | ||
| projectName | Yes | Human-readable project name to store this repository under | |
| pullRequestLimit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the write nature is partly covered structurally. The description adds real context by stating that fetched data is tagged as 'verified' evidence because it is directly observed — a behavioral trait not in the annotations. It does not, however, explain the idempotentHint=false implication (what happens on re-indexing the same repo) or auth requirements beyond 'token-accessible'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the operation and its scope, then state the evidence-classification consequence. No filler; every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should ideally describe what the call returns or confirms. For a state-mutating index operation with idempotentHint=false and no output schema, the absence of return/re-run behavior leaves a real gap, though the tool's core effect (storing verified evidence) is covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%; commitLimit and pullRequestLimit have no descriptions. The description's mention of 'recent commits' and 'pull requests' hints that those parameters bound what is retrieved, but it adds no syntax, defaults, or meaning for candidateId and projectName beyond the schema. Baseline 3 is appropriate given partial coverage and only marginal added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Indexes') and resource ('GitHub repository') and enumerates what is pulled: commits, PRs, README/architecture docs, and dependency manifests. This is clearly distinguishable from sibling tools like analyse_job_description or generate_star_answer, though it doesn't explicitly differentiate itself from careerproof_add_project_evidence, which also writes evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the operation by name but gives no explicit when-to-use/when-not guidance and names no alternatives. Given the sibling set includes careerproof_add_project_evidence, an agent has no stated rule for choosing between manual evidence entry and automated repo indexing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_match_requirementsMatch job requirements to evidenceB
Scores every requirement extracted for a job against all stored evidence (indexed repositories and/or CV claims), returning a match strength (strong / moderate / weak / unproven) plus the specific evidence cited for each requirement.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | Job ID returned by careerproof_analyse_job_description | |
| projectId | No | Restrict matching to one project | |
| candidateId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the matching scope (all requirements vs all stored evidence), the granularity (per requirement), and the returned match-strength categories plus cited evidence. It does not state that the operation is read-only, whether results are persisted, or the cost of matching against all evidence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the action and keeps the return values at the end; every clause earns its place. It borders on run-on but wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does cover the return shape, which is a plus. However, for a 3-parameter tool with an undocumented candidateId and no annotation coverage, it leaves the filtering semantics and operational side effects (persistence, read-only nature) unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 67% and the description adds no parameter meaning at all. candidateId is undocumented in both places, and 'all stored evidence' mildly obscures the projectId/candidateId restriction options rather than explaining them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: scoring every extracted job requirement against all stored evidence, and it names the output vocabulary (strong/moderate/weak/unproven). This clearly separates it from siblings like careerproof_find_evidence_gaps or careerproof_analyse_job_description, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, prerequisite, or alternative guidance. The only pipeline hint (jobId comes from careerproof_analyse_job_description) lives in the schema, not the description. An agent must infer that this is the step after job analysis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
careerproof_score_interview_answerScore a practice interview answerA
Rule-based (non-LLM) check of a candidate's practice answer: STAR structure presence, first-person action verbs, measurable results, length vs. a target word count, and keyword relevance to the target competency. Returns a score and specific, actionable feedback rather than a black-box grade.
| Name | Required | Description | Default |
|---|---|---|---|
| answerText | Yes | The candidate's practice answer | |
| competency | No | The competency the answer is meant to address | |
| questionId | No | Interview question ID to attach this score to | |
| maximumWords | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses that the check is rule-based/non-LLM, lists what is evaluated, and promises actionable feedback rather than a black-box grade. It does not clarify whether the optional questionId causes the score to be persisted or attached, leaving a side-effect gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. It front-loads the core check and follows with the return behavior, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four parameters, no output schema, and no annotations, the description provides enough context for correct invocation: it explains what is checked and the nature of the returned score and feedback. It stops short of detailing score range, feedback format, or whether the score is persisted when questionId is supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, and the description adds meaning beyond the schema by clarifying that maximumWords functions as a target word count and that competency drives keyword relevance. The answerText parameter is also directly described as the candidate's practice answer. Only questionId is left to the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it checks a candidate's practice answer, enumerating five concrete dimensions (STAR structure, action verbs, measurable results, length, keyword relevance). It is clearly distinguishable from sibling generation tools like generate_star_answer, but it does not explicitly name or differentiate them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for evaluating practice answers, but it gives no explicit when-to-use guidance, prerequisites, or alternatives (e.g., when to score versus when to generate a STAR answer). Usage is inferable from the title and first sentence, which meets the minimum viable threshold.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
careerproof_add_candidate_profile - First observed
careerproof_add_project_evidence - First observed
careerproof_analyse_job_description - First observed
careerproof_export_preparation_pack - First observed
careerproof_find_evidence - First observed
careerproof_find_evidence_gaps - First observed
careerproof_generate_interview_questions - First observed
careerproof_generate_star_answer - First observed
careerproof_index_repository - First observed
careerproof_match_requirements - First observed
careerproof_score_interview_answer
TDQS
Scored across 11 tools
Each tool has a clearly distinct role: evidence ingestion (profile, repository, manual), analysis (job description, matching, gaps, scoring), generation (STAR, questions), and export. Potential confusion between find_evidence, find_evidence_gaps, and match_requirements is resolved by descriptions specifying search vs. gap-only vs. scoring. No two tools appear interchangeable.
All 11 tools use a consistent `careerproof_` prefix followed by snake_case verb_noun construction (add_, index_, score_, export_, find_, analyse_, match_, generate_). The pattern is predictable and readable. Only minor variance is verb choice (analyse vs. analyze) but not inconsistent.
11 tools is well-scoped for a career-evidence preparation server, covering ingestion, analysis, generation, and export without obvious bloat. Each tool earns its place by handling a distinct workflow step. Falls squarely within the ideal 3-15 range.
Core lifecycle is covered: ingest evidence (CV, repo, manual), analyze job requirements, match against evidence, identify gaps, generate answers/questions, and export a pack. Minor gaps exist for updating/deleting evidence items or profiles, but these are not critical for the stated prep workflow. Agent can work around via re-ingestion or new evidence.
Maintenance
Related MCP Connectors
Generate tailored, ATS-optimized resume PDFs and cover letters from a job description, over MCP.
Job search and interview prep MCP. 15 tools, OAuth 2.1, cross-LLM. four-leaf.ai.
Screens public GitHub repos and PRs to generate risk maps, findings, and merge-readiness signals.
Job application tracker for developers - AI agents write over MCP, you review in a dashboard.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables interview preparation by analyzing resumes and job descriptions, generating role-specific questions, and evaluating answers using MCP tools integrated with an OpenAI agent.5MIT
- AlicenseNot gradedqualityDmaintenanceProvides AI coding agents with structured intelligence about any GitHub repository including overview, PRs, contributors, hot files, CI status, and dependencies via a hosted MCP endpoint.41 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables JD-aware resume matching through MCP tools, providing deterministic scoring, gap analysis, bullet rewrites, and tailored cover letter generation.308 npmMIT
- AlicenseAqualityCmaintenanceEnables MCP clients to analyze resumes for ATS compatibility, parse job descriptions, get optimization and roast-style critiques, and generate tailored resumes with shareable preview links.641 npmMIT