WebMCP
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@WebMCPFind me senior ML jobs and propose resume edits from my facts."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Tailor
An agent that can rewrite your resume, but not rewrite your experience.
WebMCP turns the resume into an evidence-bound workspace: the agent proposes, the page constrains, the human approves.
▸ webmcp-openaihac-production.up.railway.app
Open it in ChatGPT's built-in browser and the status strip reads WebMCP active; the tools register against the live origin. In an ordinary browser it is fully usable by hand and says WebMCP not detected rather than implying a tool surface that is not there.
No sign-in, no backend, no empty state — the profile, resume and 120 postings
ship in the bundle. ?reset=1
restores the demo to how you first found it.
If you only have five minutes
Three files carry the argument:
File | What it shows |
13 tools and their scopes. The registered set changes with what is actually possible — see the lifecycle section below. | |
The grounding check. Its header states what it does not do, in full. | |
36 adversarial cases trying to sneak false claims past it, including one that succeeds. |
One interaction proves it. Open a posting, ask the agent to add Kubernetes
to your skills, and watch propose_resume_edits come back refused with the
offending token named — while the same call queues a line about OpenCV, because
a fact you wrote supports that one. Then watch prepare_application disappear
from the agent's tool list until you have reviewed the diff.
Everything below is the detail behind those two sentences. The screenshots are real, captured at the width the demo is recorded at.
Related MCP server: JobFindsMe
What this is
A job board over 120 real ML postings, where the resume is an editable object and an AI agent is a first-class user of the page — working alongside the human rather than instead of them.
The agent can search postings, read the user's fact bank, and propose resume rewrites. It cannot apply them, and it cannot write a line whose claims do not trace back to something the user recorded.
What this does not do
Stated early because assuming otherwise would make everything below misleading.
Nothing is ever sent to an employer. submit_application marks a row
submitted in local state and stops. There is no employer side, no outbound
request, no integration with any ATS or job board. The confirmation modal is
real — it shows exactly what would be sent, and the human must click — but the
click ends in localStorage.
There is no account and no server. All state is local. Clear your browser
storage and you are back to the shipped demo; that is what ?reset=1 does
deliberately.
The job data is a build-time snapshot, not a live feed. 120 postings fetched
once by npm run fetch-jobs and committed. Postings age and links die. Every row
keeps its source and canonical URL so a reader can check the original.
The application tracker is a list, not a pipeline. It records what you prepared and marks what you confirmed. It does not follow up, remind, or know anything an employer did.
What is real is the constraint: the guard, the tool scoping, and the evidence behind every tag — all of which run on the data actually in this repo.
The design claim: approval changes capability, not permission
This is the part worth taking from the project.
prepare_application is not disabled while edits await review — it is not
registered. The capability does not exist.
Most human-in-the-loop designs are a permission check: the tool is there, the agent calls it, the app says no. That leaves the agent free to plan around a capability it does not have, and to argue with the refusal. WebMCP lets the page change the tool surface itself, so the constraint is expressed as absence.
The registered set mirrors what is possible right now:
7 at rest get_workspace_state · get_profile_facts · get_resume · get_applications
search_jobs · open_job · request_profile_fact
11 job open + get_job_details · get_fit_gaps · propose_resume_edits · prepare_application
11 edits pending + withdraw_edit, − prepare_application ← same count, different surface
12 ready to send + submit_application

Before and after a proposal. withdraw_edit arrives as prepare_application
leaves, struck through. The count is identical; the surface is not — which is
why the app shows names rather than a number.
Each scope is one AbortController; closing a job aborts it and unregisters
that scope in a single operation. The store-level refusal still exists for a
stale reference held from an earlier turn, but it is the second line of defence.
The first is that there is nothing to call.
Every scope change is also self-describing, because an agent that finds a tool missing has nothing to read:
1 of 1 edit queued for review. Now available: withdraw_edit. No longer available: prepare_application (needs a job open and no edits pending review).
Leading questions: three layers, each because the last one failed
Everything above is a design claim. This section is a finding — the only part of this document containing evidence that could not have been arranged, because it is a record of a real agent ignoring instructions I wrote.
The agent can ask the user to add a fact. No tool available to the agent can
write to the fact bank — that is the precise claim, and it is the one that is
tested. A writing function does exist: saveFactRequest, called when a human
clicks the confirm button in the panel. That is the design, not a loophole. What
the test establishes is that nothing the agent can reach calls it —
npm run misuse-tests fires all 13 tools with arguments designed to write a fact
and asserts the bank is byte-identical afterwards, then clicks the button and
asserts it changed.
But asking has its own failure mode, and finding it took three attempts. Each layer exists because a real ChatGPT session showed the previous one was not enough.
Layer 1 — the description. It originally read "Ask the human to add a fact… use it when get_fit_gaps reports something missing that you believe the candidate actually has."
A live session: the agent replied "Once you confirm, I'll add your TensorFlow and Kubernetes experience to the local profile fact bank."
It believed it could write. The description had taught it that. Rewritten to open with DOES NOT ADD ANYTHING and to say plainly that a missing requirement is evidence the user does not have it.
Layer 2 — the description was not enough.
Next session: the agent correctly said "the app prevents me from inserting them until they're confirmed in your profile" — then opened profile questions for TensorFlow and Kubernetes anyway, purely because
get_fit_gapslisted them missing.
It read the instruction and did it anyway. So the tool now detects when a claim names something the open posting requires and the fact bank does not support, and returns a warning that leads the summary.
Layer 3 — the agent might not relay it. Layer 2 depends on the agent's cooperation, and layer 1 already proved that is not reliable. The same warning appears on screen, where it does not depend on the agent at all:
This question came from the job posting, not from your profile. Nothing you have recorded mentions kubernetes, and this posting requires it. Add it only if you have genuinely done this work. A posting asking for something is not a reason to claim it.
The confirm button reads "Yes, I have done this", not "Add to fact bank".
It warns rather than refuses, deliberately: the app cannot hear the conversation, so it cannot tell the agent inferred this from a gap from the user just said they know Kubernetes. Refusing would block the honest case to stop the dishonest one.
What the constraint looks like
Both captured at 760px, the width the demo is recorded at.
A claim held back

The agent proposed "Backend: FastAPI, Flask, Docker, Linux, Git. Familiar with
Kubernetes and TensorFlow." citing f_backend. Two terms in that sentence trace
to nothing the user recorded, so the edit was not queued. Hedging does not help —
"familiar with" still contains the token.
A claim that passes, and what it rests on

The same call queued this one. The panel underneath is the point: the claim rests on three facts in the user's own words, quoted rather than referenced by id. The strikethrough is context; the accent-barred insertion is the proposal.
The grounding constraint
Every resume edit and the cover note pass through the same check. A line is queued only if every technology term, product name, number, superlative and seniority claim in it is grounded — present in a cited fact, or already in the block being rewritten.
Refusals name the offending tokens, so the agent can correct itself rather than retry blindly.
// refused — f_backend says nothing about either
{
"reason": "unsupported_claim",
"offendingTokens": ["tensorflow", "kubernetes"],
"hint": "Nothing in the cited facts or the original block supports: tensorflow, kubernetes.
Either cite a fact that does, call request_profile_fact to ask the user, or drop the claim."
}// queued — every technology resolves to a cited fact, both numbers appear verbatim
{
"targetBlockId": "b_exp_wagon",
"newText": "Three-stage computer vision pipeline for railway wagons: detection mAP@50 0.994
and ResNet18 recognition at 99.76% validation accuracy.",
"sourceFactIds": ["a_wagon_pipeline", "a_wagon_metrics"]
}Round 0.994 to 0.99+ and it is refused: rounding a metric invents a different
metric. "near-perfect accuracy" is refused as a superlative standing in for a
number. "Led the team" is refused as a seniority claim no fact supports and no
numeric check would catch.
npm run guard-tests runs 36 adversarial cases. 35 behave as expected, with no
false refusals among them — which matters as much as the refusals, because a
guard that blocks honest work is a guard people switch off.
That number was wrong once, and how it was wrong is the point. An earlier
version of this section claimed zero false refusals across 30 cases. Then a
recorded ChatGPT session refused "CNNs" and "APIs" — both backed by facts
the candidate wrote, both differing from the stored token only by an s. The
check compared surface strings, so CNN matched and CNNs did not. The suite
had no case of that shape, so it reported a clean sweep while the guard was
blocking true claims on camera.
Both sides are now reduced to one form before comparison — case, possessives and
punctuation are removed, and a trailing plural -s only when three or more
characters remain, which keeps AWS, GCP and RAG intact. Nothing further:
no stemming, no prefix matching, no edit distance, because any of those would let
Kubernetes reach a grounded word and that is exactly what must not happen. It
normalises to kubernete, which no fact contains, and is still refused. The same
session exposed a second case — x ray written without the hyphen was read as
the Ray framework, since the guard blocked X-ray but not the spaced form.
Every token in the fact bank was then audited against the forms a person would actually write — 684 variants across plurals, possessives, casing, hyphenation and spacing. Zero false refusals remain, and six cases of this shape are now in the suite, including two that check the normalisation did not overreach.
So: read the number as a floor, not a proof. I wrote both the guard and the cases that attack it, so the suite measures the failures I thought to look for — and a real session found a class I had not. Five of the behaviours it now catches were added after the guard was written: spelled-out numbers, superlatives standing in for metrics, unquantified seniority verbs, a narrower product name claimed from a broader concept, and this one. The suite is in the repo to be re-run and extended, and the one case it still fails is left in it.
What grounding is not
Grounding is not truth. The page checks that a claim traces to the user's own record. It cannot know whether that record is accurate. If the fact bank is wrong, the resume will be wrong, and the guard will pass it.
Grounding is not entailment. This one has a worked example, and it fails:
"Built an OCR pipeline that reads structured fields directly from X-ray scans."
cites a_ocr_declaration (OCR on scanned documents) + a_xray_loaded (X-ray classifier)Every ingredient is attested. The sentence asserts a combination neither fact supports — those two projects never met. The guard passes it. No amount of token matching will catch that; it needs entailment checking, which this is not.
Rather than pretend otherwise, the app marks any edit citing more than one distinct project and puts the decision where it belongs:
Needs your judgement. This sentence draws on 2 separate pieces of work. Each term is backed by a fact above, but nothing checks that the combination describes something that actually happened.
The failing case is kept in the test suite as an expected failure. It is surfaced, not solved.
No server-side model call
There is no backend, no API key and no LLM call in this codebase. The page exposes capability, not cognition.
That is the WebMCP-native choice rather than a shortcut. The intelligence is already in the room — it is the agent the user brought. A page that also called a model would be second-guessing it, and would need a key, a backend and a per-user cost.
It also puts the trust boundary in the right place. The page's job is not to be
smart; it is to be checkable. get_fit_gaps is a deterministic set
comparison. The guard is string matching over a controlled vocabulary. Neither
involves a model, which is exactly why the agent's claims can be checked rather
than believed.
The alias layer, and a false negative that matters
Aliases are derived from the corpus by measurement, not by intuition. The proof is what the measurement threw out — three obvious-looking aliases that a plausible guess would have accepted:
Candidate alias | What the corpus actually shows | Verdict |
| 10 hits, ~half of them "apply with your CV in English" | rejected |
| 37 hits, dominated by "serving 50,000+ customers" | rejected |
| 16 hits, ~80% text classification | profile-only |
Each was checked against the 175 fetched postings before being accepted or
dropped. classification survives only on the resume side, where it is
unambiguous; as a job matcher it would have labelled NLP roles as vision work.
The reason the layer exists at all: postings and resumes describe the same skill with different words. The corpus writes TTS, STT and ASR. This profile writes "text to speech" and "speech recognition" — forms appearing in zero of those 175 postings.
A literal comparison told a candidate with shipped speech systems that he lacked speech experience. That is the classic ATS failure, and a false negative is as damaging as a false positive: it tells someone to acquire a skill they already have, and suppresses the fact that would have supported a legitimate line.
Job tags and fact tokens now resolve through one table — 65 concepts, 207 surface forms.
ray is hyphen-guarded: the word-boundary matcher treats - as a separator, so
"X-ray" would otherwise match the Ray framework — and this profile is full of
X-ray inspection work.
Job data
120 real postings, 115 remote, 1,099 evidence-carrying tags (541 required, 141
nice-to-have, 417 unclassified). Fetched at build time from four free, no-key
public APIs using Node's built-in fetch and no dependencies:
Jobicy · Himalayas · Arbeitnow · Remotive
Every row keeps a source field and a canonical url. Descriptions are truncated
to 1,400 characters — the full text stays on the original site. We link out, we do
not republish.
npm run fetch-jobs # rewrites src/data/jobs.json and src/data/vocabulary.jsonEach tag carries a ~140-character evidence quote from the full posting, because the stored description is truncated. Without it the app could say "this job requires Kubernetes" while the requirement sat in text the user cannot see — an unverifiable claim, which is the thing this product exists to constrain. The page must not commit the error it refuses on the agent's behalf.
required is best-effort and null when undetermined, never guessed.
Unclassified tags are a third list rather than folded into "required", because
promoting them would manufacture a requirement the posting never stated.
Running it
npm install
npm run dev # http://localhost:5173
npm run buildThe app is fully usable by hand in a browser with no WebMCP support —
document.modelContext is feature-detected and the strip says "WebMCP not
detected" rather than implying a tool surface that is not there.
npm run guard-tests # 36 adversarial guard cases (needs npm run dev)
npm run misuse-tests # 48 agent-misuse cases (needs a built preview)?reset=1 restores the shipped demo data and strips itself from the URL, so a
shared machine always starts clean.
Testing it in ChatGPT's built-in browser
The strip should read WebMCP active.
"What can you do on this page?" →
get_workspace_state, 7 tools."Find remote computer-vision roles that don't want 5+ years" → the sidebar filters visibly move.
"Open the US LBM one" → the rail goes 7 → 11.
"How well do I fit?" → 19 of 29 requirements evidenced, 2 years short of 3, with the fact the years figure came from.
"Tailor my resume for it" → diffs appear with their provenance panels.
"Say I know Kubernetes" → refused, token named. If the agent offers to ask you instead, the on-screen notice should tell you the question came from the posting and not from your profile.
Accept the rest →
prepare_applicationre-registers → "Send it" → confirmation modal.
Nothing is sent without a human click, and aborting the turn closes the modal.
Known limitations
Compositional claims get through. Every ingredient attested, the combination invented. Flagged for the human, not blocked. Fixing it needs entailment checking.
The leading-question flag only fires while a job is open. It compares the claim against the open posting's requirements, so with no job open nothing is flagged. Widening it to "any tag the profile lacks" would flag legitimate bank-completion questions too.
140 KB brotli in one bundle, not split. The job snapshot is 74 KB of that. Splitting it behind a dynamic import was measured and rejected: the job list is the first paint, so the bytes are on the critical path either way, and on a bandwidth-limited link parallelism saves nothing. The gain was ~40 ms of parse time; the cost was gating tool registration on an async load, with a failure mode where a dropped fetch leaves tools registered against no data.
sourceSpread groups projects by fact-id stem (a_wagon_* → one project). A
naming convention doing semantic work. Fine for 32 facts; it would not survive a
bank ten times the size.
Descriptions cannot be unit-tested. The misuse suite calls tools directly, so it can prove the fact bank is unwritable but not that an agent will not claim it wrote to it. That gap is only closable by running real sessions — which is how all three leading-question layers were found.
Layout
docs/TOOL_CONTRACT.md the spec; code follows it, and it wins when they disagree
docs/ABSENT_TERMS.md what the fact bank deliberately does not contain, and why
docs/img/ the screenshots above
scripts/fetch-jobs.mjs build-time fetcher, no dependencies
scripts/guard-tests.mjs adversarial guard suite
scripts/agent-misuse-tests.mjs agent-shaped misuse suite
src/lib/guard.ts the grounding check
src/lib/match.ts alias resolution, filtering, fit gaps
src/store.ts one store; UI actions and tool executes call the same functions
src/tools.ts WebMCP registration and scope lifecycleIf a tool can do something no button can, that is a bug — which is why there is exactly one implementation of each action.
License
MIT — see LICENSE.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to manage job applications, documents, and contacts via the Application Tracker workspace, with read-only or write access controlled by administrator settings.
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to discover, filter, and track job openings based on the user's local resume, without uploading data to the cloud.12MIT
- FlicenseNot gradedqualityCmaintenanceEnables job-search application tracking through a private application board, managing application facts, stage history, and interview prep documents while leaving summarization and decision-making to the connected agent.
- FlicenseNot gradedqualityBmaintenanceEnables Claude Desktop to manage a job search end-to-end: find and score job listings, tailor resumes, generate application messages, and track application history, while leaving final external actions to the user.
Related MCP Connectors
Persistent docs and memory for AI agents — read, write, organize & search a shared workspace.
Coding agents build full-stack apps in persistent workspaces and share them by link.
Your company's brain for AI agents. Cited, permission-aware knowledge across every system.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ubaydullayevgiyosiddin7-eng/WebMCP-OPENAI_hac-'
If you have feedback or need assistance with the MCP directory API, please join our Discord server