Skip to main content
Glama
gzchenhao

OpenHire — Real Job Postings, Ghost Jobs Scored

OpenHire · 开聘

A job-search radar for your AI assistant — first-party listings, every posting's real age, and your résumé never touches our servers. 让 AI 助手替你盯岗的求职雷达 —— 一手职位、岗位在架时长打分,简历不经过我们的服务器。

MCP 1.0 privacy: local-first python ≥ 3.11 license: MIT 143 employers hiring OpenHire on Glama

Every number here is checkable. Open-duration report (per-employer medians, named, monthly) · numbers.json (the provenance of every figure we quote, regenerated each refresh) · raw report data. 我们引用的每个数字都能核:岗位在架时长月报 · numbers.json

Is your company in here, and you did not put it here? Your postings are in this index because your own careers page serves them publicly. We can measure how long a role has been open; we cannot see why, and a long-open role is a question, not a verdict. Claim your company — free, no payment, ever — and say why in your own words, or ask us to remove you and we will, without arguing. 贵司被收录了、而且不是贵司提交的?点这里认领,免费, 可以用自己的话解释,也可以直接要求我们移除。

三不原则 · Three things OpenHire never does

  1. 不存简历 · No résumé, ever. 简历留在你的电脑上。过网的只有一个你自己生成的匿名指纹和岗位号;协议里没有简历字段,多传一个 resume 参数会被拒绝,CI 测试钉死。 Your résumé stays on your machine. Only an anonymous, client-generated fingerprint and a job id ever cross the wire; the protocol has no résumé field, an extra resume argument is refused, and CI pins it.

  2. 不刷假单 · Never applies for you. 它只把你送到雇主自己的投递页,从不向任何 ATS 提交申请;连表单长什么样,我们的测试也只用一次普通 GET 看,到此为止。 It hands you the employer's own apply page and never submits anything to any ATS. Even our own tests only GET the form once and stop.

  3. 不追踪 · No tracking. 不埋点、不统计谁看了什么。服务器只记一条匿名授权(指纹 + 岗位号),没有你是谁;索引本身跑在你自己的机器上。 No analytics, no view counting, no telemetry. The server keeps one anonymous authorization record (fingerprint + job id) and nothing about who you are; the index itself runs on your machine.

还有一条不在这三条里,但同样锁死:排序买不到,它是(匹配度,新鲜度)的纯函数,见下文三条隐私红线。 One more that is locked the same way: ranking cannot be bought; it is a pure function of (match, freshness). See the three red lines below.

What your agent actually sees

You ask your assistant a question in plain language. It calls search_jobs, and every row comes back carrying the employer's real posting date — so the agent can reason about staleness instead of guessing.

You: Any senior Python roles that are actually still open? Skip the stale ones.

// one row from search_jobs — trimmed to the fields that matter here
{
  "title":        "Senior Python Engineer",
  "company":      "MongoDB",
  "datePosted":   "2026-03-31",   // from the employer's ATS, not a board's refreshed label
  "days_open":    166,
  "ghost_score":  0.61,           // pure f(relist_count, datePosted) — frozen by a test
  "apply_channel":"https://boards.greenhouse.io/…",   // straight to the employer
  "verified_at":  "2026-09-02T09:47:10Z"
}

Assistant: This one has been open 166 days with a ghost_score of 0.61 — I'd deprioritise it. Here are four posted in the last three weeks instead…

ghost_score measures how long a posting has been open, not whether the employer still intends to hire. A long-open role can equally mean "hard to fill". Treat it as a reason to ask, not a verdict.


An MCP server that turns any MCP-speaking assistant — Claude, Cursor, Windsurf, Cline, ChatGPT via connectors — into a private radar for AI / Infra, autonomous-driving and embodied-AI jobs — pulled straight from 143 employers' own career sites and public ATS APIs (Greenhouse / Lever / Ashby / 北森 Beisen / Moka, plus first-party employer career sites such as Li Auto's), across the US, Europe and China (Waymo, Figure, Zoox — and Unitree, XPeng, UBTECH, Mech-Mind…). No account. No signup. No résumé upload. Ever.

Three things a job board won't do for you:

  • Surfaces how long each role has really been open. Every listing carries a ghost_score aged off the employer's real posting date where the employer's system reports one (otherwise the day this index first saw it, flagged date_signal): the "2 days ago" a board shows you can be 300 days old in the ATS.

  • Structural privacy, not a pinky-promise. There is no résumé field in the protocol; a CI test fails the build if anyone adds one. Matching runs on your machine — only an anonymous fingerprint reaches the server.

  • Ranking you can't buy. Order is a locked pure function of (match, freshness). No sponsored slots, no bidding — the signature is frozen by a test.

This is the 「哨兵 / Sentinel」 reference implementation — see design_handoff_openhire_v01/README.md for the full protocol spec.


Related MCP server: Job Application MCP

Quickstart — under a minute

Nothing to install, and no crawl to sit through. Add this to your MCP client's config:

{ "mcpServers": { "openhire": { "command": "uvx", "args": ["openhire@latest", "serve"] } } }

Then ask your assistant for a job. That is the whole setup. The server downloads the ~30 MB public index by itself in the background on first start, so searches fill in within a couple of minutes while you are already talking to it. No account, no signup, no résumé upload.

The one prerequisite is uv (it provides uvx). Without it the config above fails with nothing but "server failed to start", so install it first:

curl -LsSf https://astral.sh/uv/install.sh | sh     # macOS / Linux
irm https://astral.sh/uv/install.ps1 | iex          # Windows PowerShell

On Claude Desktop you can skip even that — download openhire-0.6.8.mcpb and double-click it. No terminal, no Python, no uv.

On Cursor, one click (it still needs uv on your PATH): Install MCP Server in Cursor

Per-client config paths and the trade-offs of uvx vs a one-time install are in Works with below.

pipx install openhire      # keeps it isolated and puts `ohp` on your PATH
ohp bootstrap              # 143 employers · ~17k live postings · no account

ohp search --required-skills rust,k8s --remote --role-family engineering
ohp search --currency CNY --role-family engineering   # e.g. CN 智驾 / robotics roles

ohp bootstrap downloads the public snapshot (~30 MB, no account) and then runs one incremental crawl to refresh verified_at and catch delistings. The crawl is the slow part — a line per employer, 20+ minutes on a cold index — and you can stop it once the snapshot is in; the index is already usable, just verified as of the last weekly refresh rather than today. The MCP path above needs none of this: serve fetches the snapshot by itself in the background.


Works with

All clients use the same MCP entry. The config below works in every MCP client and pulls the package on demand — but it does need uv present first.

New to MCP? Two shortcuts before the config below. Claude Desktop: download openhire-0.6.8.mcpb and double-click it. No terminal, no Python. Cursor — Install MCP Server in Cursor (needs uv installed). Cursor / Claude Code — or paste this to your agent: "Install the MCP server at github.com/gzchenhao/openhire. Install uv first if it is missing, then add uvx openhire@latest serve to my MCP config and tell me which file you changed."

Prerequisite: uvx ships with uv. Without it the config below fails with nothing but "server failed to start" in your client — install uv first:

curl -LsSf https://astral.sh/uv/install.sh | sh     # macOS / Linux
irm https://astral.sh/uv/install.ps1 | iex          # Windows PowerShell
{ "mcpServers": { "openhire": { "command": "uvx", "args": ["openhire@latest", "serve"] } } }

First start downloads a ~30 MB index in the background; searches fill in within a few minutes.

Stuck? Run ohp doctor. It checks the three things that all look identical from the chat window — uv missing, no index yet, server configured but not enabled — and reads every client config it can find. It runs in your terminal, which matters: if the client never started our server, nothing we wrote inside it can reach you.

Editing the config may not be the last step. Several clients require you to enable or trust a newly added server before its tools load — the config is saved, the server never starts, and the only symptom is that your assistant does not seem to know about the tools. If a search does nothing, open your client's MCP/connectors panel and check that openhire is listed and switched on. (Claude Desktop needs a full quit and reopen; Cursor and Windsurf pick it up on reload; some clients show a per-server toggle.) 改完配置不一定就完事:部分客户端需要你在设置里手动「信任 / 启用」这个 server, 工具才会加载。症状是配置明明在、助手却完全不知道有这些工具。

What uvx costs you, every time. uvx resolves the package on each invocation — measured at 7–8 s per call even with a warm cache. That is paid on every MCP session start and every CLI command. It buys you never having to manage an install. If you would rather pay once:

pipx install openhire     # then use "command": "ohp", "args": ["serve"] — process-start latency

@latest also means your tool surface can change under you without warning. Pin it when that matters: "args": ["openhire==0.6.8", "serve"].

The server auto-downloads the public job snapshot on first run if the index is empty, so ohp bootstrap is optional. If you ran pipx install openhire, "command": "ohp" works too.

Claude Desktop — %APPDATA%\Claude\claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/); quit & reopen after editing:

{ "mcpServers": { "openhire": { "command": "ohp", "args": ["serve"] } } }

Cursor — ~/.cursor/mcp.json (or a project .cursor/mcp.json):

{ "mcpServers": { "openhire": { "command": "uvx", "args": ["openhire", "serve"] } } }

Windsurf — ~/.codeium/windsurf/mcp_config.json:

{ "mcpServers": { "openhire": { "command": "uvx", "args": ["openhire", "serve"] } } }

Claude Code — this repo is also a plugin marketplace, so two commands inside Claude Code do the whole setup (still needs uv on PATH):

/plugin marketplace add gzchenhao/openhire
/plugin install openhire@openhire

Cursor plugin / Agent Plugins — the repo carries .cursor-plugin/plugin.json and the vendor-neutral plugin.json + mcp.json (agent-plugins.org), so any client that loads Agent Plugins can install it from the repo URL.

TRAE(字节) — one-click import links (TRAE asks you to confirm the config before adding it): TraeCode 国内版 · TRAE international. Or paste the uvx config above into TRAE's MCP settings by hand.

腾讯 WorkBuddy — ~/.workbuddy/mcp.json, same mcpServers shape. WorkBuddy reads the file when it starts, so after editing it quit and reopen the app (or add the server through its MCP panel's 「添加 MCP」 instead); a hand-written entry does not appear in the list until then. No uv on the machine? pip install openhire into any Python ≥ 3.11 environment and point command at that environment's ohp executable with "args": ["serve"].

First start downloads the ~30 MB public snapshot (jobs/companies only) — give it a moment. To refresh later run ohp bootstrap --force or ohp ingest. On Windows Claude Desktop from the Microsoft Store, the config is under …\Packages\<Claude package>\LocalCache\Roaming\Claude\.

Hosted / remote: ohp serve --transport streamable-http --host 0.0.0.0 --port 8000 exposes http://host:8000/mcp (also --transport sse). A Dockerfile is included.

GitHub unreachable from your network (mainland China without a proxy, some sandboxes)? The snapshot is a GitHub Release asset and the download now resumes and retries, but a blocked host stays blocked. Fetch openhire-index.db.gz through a mirror you trust or on another machine, then ohp bootstrap --snapshot-url <url-or-local-path>, or set OPENHIRE_SNAPSHOT_URL in the server's env so the auto-download uses it. ohp bootstrap --fresh skips the snapshot entirely and crawls the employers' own ATS (20+ minutes, no GitHub involved). An empty index caused by a failed download says so in the tool result (bootstrap_error) instead of looking like "no matches". 国内网络 GitHub 不通:经你信任的镜像拿到快照后 ohp bootstrap --snapshot-url <地址或本地文件>, 或在 MCP 配置的 env 里设 OPENHIRE_SNAPSHOT_URL;ohp bootstrap --fresh 不经 GitHub 直接抓。


What it does

Tool

What it gives you

search_jobs

Hard-filter the live index; every result carries verified_at, datePosted, days_open, ghost_score, remote_scope, eligible_regions, apply_channel. Filter by required_skills (AND), role_family, remote_scope, min_salary + currency, company, location (bilingual city aliases: 香港 = Hong Kong, 广州 reaches 天河区) and title (words in the job title, e.g. 招聘 / recruit — the filter for HR, finance and other roles no skill tag names).

watch_intent

Register a standing intent once — new matching jobs are waiting next time you check, even after you close the terminal. Accepts required_skills / role_family / location / title so sales / solutions roles stay out.

check_watches

Pull the matches that are new since your last check (client-pull; stdio has no push).

authorize_application

One explicit confirmation per job. It records your authorization and returns the employer's own application URL — you apply as yourself. It cannot accept a résumé.

get_company_info

Aggregate, anonymous trust signals for one employer (ghost_score_avg, active_jobs, index_built_at). Never any candidate data.

Optional, entirely local: ohp init --scan <dir> derives a skill fingerprint from your own repos. You never write a résumé; the code never leaves your machine — only an anonymous vector does.

For employers: claim your tenant

If your company is in this index, the listings came from your own public ATS — we did not ask, because we did not need to. What we cannot know is your side of it: whether a role is an evergreen talent pool rather than a stale req, or how fast you actually reply.

Claim it (中文表单). Free, verified by corporate identity — a GitHub org membership or a reply from a corporate domain — and never by payment. We answer within 3 business days, claiming leads to no paid follow-up of any kind, and your proof is used to verify and then nothing else.

No GitHub account? Email gdchenhao@qq.com with "Employer claim" and your company name in the subject — sending from your corporate domain is itself the verification — or have anyone file the form on your behalf, since what we verify is the company and not the filer. 没有 GitHub 账号?直接发邮件到 gdchenhao@qq.com,用贵司企业邮箱发出来即完成身份核验。

A claim gets you:

  • Your own note, attributed to you, beside the roles you name. We can see how long a role has been open; we cannot see why. An evergreen talent pool, a genuinely hard-to-fill role, and a neglected one look identical from outside, and only you can tell them apart.

  • evergreen / hard to fill / closed status on specific titles. A closed role stays visible while your ATS still serves it — we do not hide what your own site returns — but it is marked closed on your word so nobody else applies.

  • A correction if your ATS's date field means "requisition opened", not "went live". That one misreading makes every role you have look years old, and it is not your fault.

  • response_sla_days on every one of your postings, including ones you post later

  • claimed: true on your company, with the date

It does not get you rank, and it does not lower your ghost_score. Ordering is a locked pure function of (match, freshness) and the score is a pure function of (relist count, posting age); tests freeze both and assert the claim path touches neither. Your note sits beside the score and explains it. We would rather show a high score with your explanation than a quiet score somebody paid for.

Asking to be removed entirely is also fine, and we will not argue about it.

Verified claims live in src/openhire/seed/claims.py — in the repo, not in a private database — so each one is a reviewable diff, and the weekly rebuild re-applies them instead of quietly dropping them.

Keeping it current

The index refreshes weekly, so a search can be up to seven days behind. When the user is about to act on one employer and wants today's truth:

ohp refresh unitree          # ~1 minute · at most one crawl per employer per 6 hours

Over MCP this is the refresh_index tool. Three rules are built in, not advisory:

  • One employer per call. A full crawl is 20+ minutes and no client waits that long. An ambiguous word (robot → 11 matches) is refused with the candidate list, never fanned out.

  • Six-hour throttle per employer, checked before any network call, so a too-soon request costs the ATS nothing and returns last_refreshed_at instead of an error.

  • Not for speculative or looped calls. Each one hits somebody else's public endpoint.

That last point is the whole design constraint: our own crawl is a weekly batch we control, and handing refresh to callers turns it into our users hitting their endpoint on our behalf. The throttle is what keeps that boundary ours to keep rather than ours to spend.

Reading the three dates

Every row carries three timestamps that answer three different questions. Read together they separate an abandoned requisition from one somebody is still tending:

field

question it answers

verified_at

did the employer's ATS still return this the last time we looked?

datePosted / days_open

how long has it been open? ghost_score ages off this. The employer's own date where its system reports one; otherwise the day this index first saw the posting, and the row carries date_signal: "not_reported_by_ats" (Li Auto's first-party mirror reports no date), so read it as a lower bound

updated_at / days_since_update

when did the employer last touch it? Null when their ATS does not report one (Ashby, Lever and Beisen do not; Greenhouse and Moka do) — read null as "unknown", never as "abandoned"

ghost_score = 1.0 alone is not a verdict. Open 367 days and untouched for 367 days reads as abandoned; open 327 days but touched 13 days ago reads as a tended evergreen req. Among the rows where the ATS actually reports a last-touched date, 67% of ghost_score >= 0.99 postings were touched by the employer within the last 30 days (measured 2026-09-19; the live figure is pct_ghost_hi_touched_within_30d in docs/numbers.json, which is regenerated every refresh). ghost_reason spells out which input drove the score ("age only: open 367d, never relisted").

None of these measure intent. A long-open role can equally mean hard-to-fill — treat the numbers as a reason to ask, not a verdict.

Asking about one employer

search_jobs(company=...) takes whatever the user actually said — an id (unitree), or any part of the name in either language (宇树, Unitree, XPeng). A name this index does not carry comes back as the empty-result object naming the company, never as an unfiltered search.

ohp search --company 宇树 --role-family engineering
ohp search --company waymo --distinct          # one row per role, not one per city

--distinct (collapse_role_group over MCP) keeps one row per role_group and adds role_group_size. About 20% of a page is the same role listed once per city; folding is opt-in because each city row has its own job_id and apply_channel, which matters under a location or visa constraint.

The five protocol fields

Every listing is valid schema.org/JobPosting, plus:

  • verified_at — last moment confirmed live on the employer's own site

  • source — employer_site | ats_public_api (never a job board)

  • ghost_score — 0–1 listing-activity signal, aged off the real posting date (lower = fresher). A noise filter, not an accusation: long-open listings are often evergreen talent pools or slow pipelines — the score simply lets agents down-rank low-activity noise

  • response_sla_days — the employer's OWN committed reply window. Null on almost every row, and that null is meaningful: it is set only when an employer claims their tenant. We never infer or estimate it, because we cannot observe a reply even in principle — the application deep-links to the employer and never touches this server. Read null as "no employer has claimed this tenant", not as "missing" or "slow". Employers: claim yours — free, verified by corporate identity, never by payment. It buys a verified badge, the ability to correct listing status (an evergreen pool carrying an unfair staleness score, say), and this field. It does not buy rank: ranking is a locked pure function of (match, freshness), and a test freezes that signature.

  • apply_channel — always the employer's own application URL, deep-linked to the specific job

Privacy Policy

Short version: there is no résumé field in the protocol, matching runs on your machine, and the only user-originated value the server ever stores is an anonymous client-generated fingerprint. No analytics, no telemetry, no third-party sharing. Full policy: docs/PRIVACY.md.

Privacy model

Résumé / PII upload

never — matching runs locally; a résumé never transits the server, and we never store one

What the server sees

one anonymous, client-generated fingerprint + hard filters

Repo scan

local-only · personal projects · explicit consent · opt-out anytime

Job sources

first-party only: employer career pages + public ATS APIs (Greenhouse / Lever / Ashby)

First-run data — the snapshot vs. fresh

ohp bootstrap (default) downloads a small public index snapshot (a GitHub Release asset — companies + jobs only, zero user data) and then runs one incremental crawl to refresh verified_at / delisting. --fresh skips the snapshot and crawls the public ATS from scratch with the free offline heuristic extractor. Either way: no account, no PII.

Two things that surprise people:

  • The incremental crawl is slow and quiet. On a cold index it can run for 20+ minutes with no output. It is working, not hung. If you only want the data, ohp serve skips it entirely — the server downloads the snapshot on first start and is answering in seconds.

  • The snapshot URL is pinned to the v0.1.0 tag on purpose. It looks stale; it is not. That asset is overwritten in place every Monday by a scheduled workflow, so the URL is a stable address for always-current data. Pinning it to the newest tag would break every client the moment a release is cut.

Three rules this project will never break

  1. Your résumé stays on your machine — it never transits the server, and we never store it.

  2. Ranking is not for sale — it is only f(match_quality, freshness), a locked pure function.

  3. Employers pay only for authorized, delivered outcomes — never for exposure. (v0.1 has no billing at all.)

These are enforced by CI (tests/test_privacy.py, tests/test_ranking.py, tests/test_snapshot.py).

Development

python -m venv .venv && . .venv/Scripts/activate   # Windows
pip install -e ".[dev]"
pytest        # privacy red lines + ranking + snapshot must be green

Set OPENHIRE_DATABASE_URL=postgresql+psycopg://… to run against Postgres instead of the default local SQLite file (~/.openhire/openhire.db).

Roadmap

  • v0.2 – v0.3 (shipped) — CN ATS adapters (北森 Beisen + Moka) · weekly auto-refreshed public snapshot · ghost_score public beta · 143 employers across US / EU / China

  • next — Employer claim + verified badges — employers can reserve their claim today via a corporate-identity GitHub issue (zero-cost now; badges + listing-status control ship next) · response-SLA enforcement (7-day auto-delist) · redacted proof-of-fit — an anonymous, candidate-authorized match summary that travels with an application (skills overlap only; identity never included, résumés still never transit the server)

  • v1.0 — Open, vendor-neutral schema extension for AI-readable job postings

FAQ

Where does the job data come from? Directly from 143 employers' own public ATS APIs (Greenhouse, Lever, Ashby, 北森 Beisen, Moka, plus first-party employer career sites such as Li Auto's), the same endpoints that power their careers pages. No third-party job boards. source is always ats_public_api, and verified_at records the last time we confirmed each posting live. The public index is auto-refreshed weekly, so a fresh ohp bootstrap starts from recent data.

Why should I trust ghost_score? It's a pure, open, unpurchasable function — min(1, 0.15·relist_count + staleness) aged off the real ATS posting date, not our crawl date. The formula lives in pipeline/ghost_score.py, is unit-tested, and takes no money as input (red line #2). Long-open, repeatedly-relisted postings score higher; you can always re-rank client-side. Read it as signal-to-noise, not bad faith: plenty of high-scoring listings are legitimate evergreen talent pools. Employers who want their listing activity represented accurately can claim their tenant (see Roadmap).

Does my résumé actually go through the server — really? No. There is no résumé anywhere in the protocol. authorize_application has no résumé/file parameter (it structurally cannot accept one), matching runs on your machine, and the only thing that ever transits the server is an anonymous fingerprint the client generates, like #a3f9-k2p7-x8q1. This is enforced by tests/test_privacy.py, and the published snapshot carries zero user data (tests/test_snapshot.py).

Does it support China (中国区)? Yes — this is what sets OpenHire apart. Employers on 北森 Beisen (<tenant>.zhiye.com) and Moka (app.mokahr.com) are indexed: 20+ autonomous-driving / robotics / embodied-AI companies including 宇树 Unitree, 小鹏 XPeng, 优必选 UBTECH, 梅卡曼德 Mech-Mind, 速腾聚创 RoboSense, 元戎启行 DeepRoute, 星海图 Galaxea, 傅利叶 Fourier, 普渡 Pudu. Pay published as 月薪 keeps its real period (salary_period), so a salary floor no longer silently drops Chinese roles.

飞书招聘 (Feishu Hire) is not supported and won't be: it signs its job-list requests with a ByteDance _signature and gates them behind a captcha SDK, so its listings are not publicly readable. We don't break anti-bot measures.

How do I get a company added? Open a Company inclusion request issue (title it with the company + its ATS URL) — this is the best way to contribute. If you code, add it to src/openhire/seed/candidates.py (company slug + ATS vendor/tenant) and open a PR; the seeder validates tenants against the live API.

License

MIT © OpenHire Protocol · PRs welcome.


Built by a deep-tech headhunter who does not write code, pair-programming with Claude Code. Full acceptance reports, including the mistakes, in reports/.

Available Tools

6 tools
authorize_applicationAuthorize applicationA

Record an authorized, employer-direct application. REFUSES résumés.

(Formerly apply — renamed to make explicit that this only records the user's authorization to apply as themselves; it never submits anything on their behalf.)

This tool never accepts a résumé, file, cover letter, name, email or phone — a résumé never transits the server. It only takes a job_id, an anonymous fingerprint, and an explicit per-job authorization. On success it returns the apply_channel (the employer's own application URL) for the user to submit as themselves, plus resume_transmitted=false. Do NOT paste résumé content into any argument.

REFUSES means refuses: any argument this tool does not declare (a resume key, a cv, a file, anything) is answered with ERR_PII_NOT_ACCEPTED and nothing is recorded. It is not silently dropped, so a client that sends one finds out.

Args: job_id: the job to apply to (from search_jobs / check_watches). fingerprint: the user's anonymous fingerprint. authorized: must be true — explicit per-job consent.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
authorizedYes
fingerprintYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond annotations: it discloses that undeclared arguments are refused with ERR_PII_NOT_ACCEPTED, nothing is recorded on refusal, a résumé never transits the server, and success returns resume_transmitted=false. These are critical behavioral details for an AI agent and are not visible from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior and refusal rule are front-loaded, and the structure uses bold and line breaks effectively. Some redundancy exists around the résumé refusal and the rename history, but the emphasis is justified for a privacy-sensitive tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a side-effectful tool with no output schema, the description covers success returns, refusal behavior, parameter sources, and PII constraints. An agent has enough information to invoke it correctly and avoid dangerous misuse.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description compensates by explaining job_id as coming from search_jobs/check_watches, fingerprint as anonymous, and authorized as a mandatory true consent flag. It could add more detail on how to obtain or format the fingerprint, but it covers the essential meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it records an authorized, employer-direct application and explicitly says it never submits anything on the user's behalf. This clearly distinguishes it from submission-like actions and from siblings such as watch_intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear behavioral boundaries: use it only to record authorization, never to attach files, and direct the user to submit via the returned apply_channel. It does not explicitly name sibling tools or state when not to use it, but the context is strong enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_watchesCheck watchesA

Pull matches that are new since this fingerprint's last check.

stdio has no server push, so clients pull: call this at the start of a session. Returns the new matches per watch and advances each watch's last-notified marker.

The FIRST pull on a watch returns everything matching it, not an increment: nothing has been reported for that watch before, so the whole standing set is new to the user. Each result says which it is via is_first_pull; do not present a first pull to the user as "postings that appeared since last time".

Each result also carries total_matching and truncated: a broad watch can match hundreds of postings and only the 100 best-ranked are returned. Tell the user when truncated is true; the rest is not re-reported later.

ParametersJSON Schema
NameRequiredDescriptionDefault
fingerprintYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the sparse annotations, the description reveals critical behavioral traits: it 'advances each watch's last-notified marker' (a side effect), a first pull returns the entire standing set rather than an increment, and results can be truncated to 100 with the remainder not re-reported later. These disclosures go far beyond the annotations and prevent misinterpretation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action, then adds three distinct behavioral nuances (side effect, first-pull semantics, truncation) in separate paragraphs. Every sentence earns its place, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema, the description specifies the key result fields (is_first_pull, total_matching, truncated) and explains the edge cases that matter for correct interpretation. It also covers when to call the tool and what the side effect is, making it self-sufficient for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description carries the parameter explanation. It clarifies that fingerprint is tied to a last-check state and that results are per watch, adding meaning the bare string schema lacks. However, it leaves some ambiguity about what the fingerprint identifies (a user, device, or watch set), so it is not fully explicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Pull matches that are new since this fingerprint's last check.' This clearly identifies the tool as a polling operation for watch results and distinguishes it from sibling tools like search_jobs (search) and watch_intent (watch creation), even without naming them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a strong when-to-use directive: 'call this at the start of a session' and explains the push/pull rationale. However, it does not explicitly contrast with alternatives (e.g., when to use search_jobs instead), so it stops short of the full when-not-to-use guidance needed for a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_company_infoCompany trust signalsA
Read-onlyIdempotent

Aggregate, anonymous trust signals for one employer.

Returns ghost_score_avg, active_jobs, and index_built_at (when the index was last built). NEVER returns any individual candidate data: the server holds none.

posting_dates_reported is false when no live posting of this employer carries the employer's own posting date (first-party mirrors such as Li Auto report none), and postings_without_reported_date counts those rows; their days_open, and so this employer's median_days_open and ghost_score_avg, are counted from the day this index first saw each posting, a lower bound on age rather than the employer's timeline.

claimed is true only when the employer has claimed this tenant and we verified them by corporate identity, never by payment, and it never affects ranking. It is the only signal here that comes from the employer rather than from their public ATS data, and it is what makes response_sla_days non-null on their postings.

Takes an id or any part of the name in either language, like search_jobs' company.

Known-but-not-indexed employers (Momenta, 小马智行 Pony.ai, 智元 AgiBot, MiniMax, 智谱 Zhipu, 商汤 SenseTime, 逐际动力 LimX, 自变量 X Square, 千寻智能 Spirit AI, 加速进化 Booster) return a structured answer instead of ERR_COMPANY_NOT_FOUND: indexed: false, careers_url (their own portal), reason (their careers site runs on Feishu Recruitment, whose job-list API requires a request signature; we treat that as access control and do not work around it) and employer_opt_in (the employer can authorize the read-only Feishu open-platform scopes hire:site:readonly and hire:site_job_post:readonly). 蔚来 NIO gets the same shape with its own reason: its careers page's security policy blocks this crawler (HTTP 567), which we do not work around. There are no trust signals in that answer because we hold none of their postings; do not read the absence as a verdict on the employer.

ParametersJSON Schema
NameRequiredDescriptionDefault
company_idYes

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses individual-candidate privacy, that posting ages are lower-bound estimates from first index sighting, that claimed is verified by corporate identity not payment and never affects ranking, and that non-indexed employers deliberately return no signals to avoid false verdicts. This is exceptionally rich behavioral context and contains no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with the core purpose and return values, then parcels out important caveats in separate paragraphs. Each paragraph earns its place for a tool with this much edge-case nuance, though a tighter phrasing of the known-but-not-indexed block would improve scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one required parameter, no output schema, and nuanced trust signals, the description covers the main return fields, explains the semantics of every unusual field, documents the non-indexed response shape, and provides a caveat against misreading missing signals. An agent can call this tool and interpret its result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the full burden for company_id. It explains that the parameter accepts an id or any part of the name in either language, matching search_jobs' company, which meaningfully exceeds the bare schema. It stops short of giving precise resolution rules, but the guidance is usable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states the verb and resource clearly: aggregate anonymous trust signals for one employer, and the return list plus edge-case sentences make the scope concrete. It does not explicitly distinguish itself from siblings like search_jobs, though the employer-signal focus makes the difference easy to infer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for employer trust-signal lookups and explains behavior for non-indexed employers, but it never explicitly says when to choose it over sibling tools or when not to use it. No alternative is named or excluded, so usage guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refresh_indexRefresh one employerA

Re-crawl ONE employer's public ATS now. Takes about a minute. Throttled to 6h.

The index is refreshed weekly, so search_jobs can be up to seven days behind and check_watches has nothing new to report until it moves. This is the manual nudge for the case that matters: the user is about to act on one employer and wants today's truth.

Rules worth knowing before you call it:

  • ONE employer per call. A full crawl is 20+ minutes and no client will wait; a vague word like "robot" is refused with the list of candidates rather than fanned out into eleven live crawls.

  • At most one crawl per employer per 6 hours. A throttled call returns immediately with refreshed: false, reason: "throttled" and last_refreshed_at — no network request is made. That is not an error: it means the data you already hold is that fresh.

  • An employer we know but deliberately do not index (the Feishu-hosted ones and 蔚来 NIO, see search_jobs) returns reason: "known_not_indexed" with the same portal, reason and employer_opt_in the other tools give, never unknown_company.

  • Do NOT call this speculatively or in a loop. Every call hits somebody else's public endpoint. Search first; refresh only when the user needs today's state of one employer.

Args: company: one employer — an id ("unitree") or any part of the name in either language ("宇树", "XPeng"). Ambiguous input is refused, not guessed.

Returns: refreshed plus last_refreshed_at / next_allowed_at; when it did run, also jobs_new / jobs_updated / jobs_delisted / jobs_unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
companyYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as non-read-only and non-idempotent, but the description adds substantial behavioral detail: it hits an external public endpoint, takes about a minute, is throttled to one crawl per employer per 6 hours, returns immediately with refreshed:false when throttled, and handles known-not-indexed employers. These are exactly the side effects and constraints an agent needs to know before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core behavior and duration, then organized into scannable rule bullets. Every sentence adds operational value—rate limits, failure modes, and expected return fields are all relevant. Despite its length, there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema, the description is remarkably complete. It covers return fields, throttling behavior, known-not-indexed cases, ambiguous input handling, and the operational context around the weekly refresh. An agent has everything needed to decide when to call and what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only a bare 'company' string with no description, and schema coverage is 0%, so the description carries the full burden. It explains that company can be an id or any part of the name in either language, gives concrete examples, and states that ambiguous input is refused rather than guessed. This fully compensates for the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: re-crawl ONE employer's public ATS. It clearly distinguishes itself from search_jobs and check_watches by explaining the weekly refresh lag and the manual nudge use case. The scope is unambiguous—one employer per call, not a broad index operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: when the user is about to act on one employer and needs today's truth. It also gives strong when-not-to-use guidance, including not to call speculatively or in a loop, and instructs to search first. This is exemplary routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_jobsSearch jobsA
Read-onlyIdempotent

Search the live job index by hard filters; returns ranked JobPosting[].

The server does ONLY a hard filter plus a fixed ranking of match-quality × freshness — precise re-ranking is left to you, the client, which holds the user's context. Every result includes the five protocol fields (verified_at, source, ghost_score, response_sla_days, apply_channel) plus datePosted, days_open, remote_scope, eligible_regions, role_group and ghost_reason.

verified_at and ghost_score answer DIFFERENT questions and routinely disagree: a posting confirmed live today can score 1.0. Live means the employer's ATS still returns it; the score means it has been returned for a long time, or keeps being relisted. ghost_reason says which input drove the score ("age only: open 367d, never relisted") so you can tell the user that instead of a bare number. Neither field measures intent — a long-open role can equally mean hard-to-fill.

How to say it: ghost_score is a measurement of time on the market, never a verdict that a posting is fake. Do not call a posting a ghost, zombie, fake job or 僵尸岗 on the strength of it, and do not tell the user "don't apply" because of it. Report what was measured — "open 1,616 days, never relisted, the ATS reports no last-touched date" — and let the user weigh it. A 1.0 can be an abandoned req or a role that has been genuinely hard to fill for four years; the number cannot tell those apart.

response_sla_days is null on almost every row, and that is a meaningful null: it is the employer's OWN committed reply window, set only when they claim their tenant. We never infer or estimate it — the application deep-links to the employer and never touches this server, so we cannot observe a reply even in principle. Read null as "no employer has claimed this tenant", never as "missing data" or "slow".

employer_correction appears only when the employer has claimed this tenant and said something about this specific role. It is the employer's own account, verified by corporate identity, and it sits BESIDE ghost_score, never on it: status: "evergreen" explains a high score (they hire continuously, so there is no single opening to fill), it does not lower it. status: "closed" means they say they are no longer hiring even though their ATS still returns the row, so stop sending people there. date_semantics: "requisition_created" means their ATS reports the day the req was opened internally rather than the day it went live, so days_open overstates for that employer. Treat all of it as the employer's claim, attributed, not as our measurement.

days_since_update is the third date and the one that usually settles it: the employer's own last-touched timestamp from their ATS. Two rows can both score 1.0 and mean opposite things — open 367d and untouched for 367d reads as abandoned; open 327d but touched 13 days ago reads as a tended evergreen req. Among rows where the ATS reports one at all, the share of ghost>=0.99 postings touched by the employer inside 30 days has ranged from about a third to three quarters across refreshes; read the current value from pct_ghost_hi_touched_within_30d in docs/numbers.json rather than quoting a figure.

Null is NOT "abandoned": Ashby, Lever and Beisen do not report a last-touched date, and those rows carry update_signal: "not_reported_by_ats" instead. A row carrying date_signal: "not_reported_by_ats" goes one step further: its source reports no posting date at all (Li Auto's first-party mirror is one), so its datePosted and days_open are counted from the day this index first saw it, not from the employer's own date. Read those as a lower bound on age, never as the employer's timeline. Treating that null as "nobody has touched this in a year" would describe the vendor's API, not the employer. Honest limit even when present: an ATS bumps it on any edit or re-publish, so it means "touched", not necessarily "content changed".

To answer "what is hiring?", pass company — do not filter client-side.

Skills are alias-aware on the request side: 感知 finds rows tagged perception, 占用网络 finds occ / occupancy / occupancy networks. If a requested tag matches NOTHING in the index, the response switches to {results, unknown_skills, suggestions} so you can tell the user which word was dropped instead of presenting a half-match as a match.

role_family filters exclude rows KNOWN to be another family; rows not yet classified (role_family null, typically the newest postings) are included, not hidden. remote_scope can also be "unknown": remote with a location we could not read; it is never reported as "worldwide" any more.

Two things worth knowing before you spend your budget:

  • role_group is shared by the same role posted in several cities — one employer may list one job 22 times, once per location. Those rows are genuinely distinct (each has its own job_id and apply_channel, which matters when the user has a location or visa constraint), but if you only need distinct opportunities, group by role_group and keep one per group. Measured, about 20% of a page is same-role repeats.

  • offset pages through the ranked list. After collapsing by role_group, call again with offset += limit to get more distinct roles. Fewer rows than limit means you reached the end. There is no server-side cursor to keep alive.

  • One call returns at most 100 rows. Ask for more and you get an object with results and truncated, not a silent first page: 100 of 198 looked exactly like "this employer has 100 jobs". truncated: true means a full page came back and there may be more, with the offset to call next in hint; truncated: false means fewer rows than the page size matched and there is no next page. For one employer, get_company_info's active_jobs is the true total.

  • Skill tags match separator-insensitively: "computer vision", "computer-vision" and "computer_vision" are one skill. The extractor emits all three spellings, so an exact-string search reached as little as 42% of the rows that had the skill.

  • An empty search does NOT return a bare list. It returns an object with results: [] plus hint, unknown_skills and suggestions, because [] alone cannot tell you whether you mistyped a tag or the market is genuinely dry. Read unknown_skills: if it is non-empty those tags exist nowhere in the index and you should retry with a suggestion; if it is empty your tags were fine and you should loosen a filter.

Args: skills: skill tags, ANY-overlap match (union), e.g. ["rust", "k8s"]. required_skills: skills that must ALL be present (AND), e.g. ["rust"]. remote: if true, only fully-remote roles. remote_scope: filter remote roles by reach: "worldwide" | "unknown" | "region_locked" | "country_locked". min_salary: salary floor. By default roles with NO stated pay are KEPT (they can't be ruled out); set require_stated_salary=true to drop them. currency: restrict to a stated-pay currency, e.g. "USD" (implies stated pay). require_stated_salary: if true, drop roles that publish no salary. role_family: coarse family filter, one of engineering | data | product | design | marketing | sales | ops | other. Any other value is refused (ERR_UNKNOWN_ROLE_FAMILY) rather than ignored: "recruiting" used to pass through and return the whole index. There is NO hr / recruiting / people family — those roles are filed under ops; use title to isolate them. Populated for most live rows, so this is an effective way to keep sales / solutions-architect roles out of an engineering search. title: caseless substrings over the job title, ANY-of, e.g. ["recruit", "招聘"] or ["感知"]. This is the filter for roles no skill tag or family can isolate: HR, recruiting, finance, legal, a specific team name. An ASCII term must start a word ("hr" reaches HR, HRBP and HR Business Partner but not Chrome); a CJK term is a plain substring. One synonym group is expanded for you: any of hr / hrbp / human resources / recruit / talent acquisition / sourcer / people ops / 人力 / 人事 / 招聘 reaches all the others, so title=["招聘"] also finds an English "Senior Technical Recruiter". Combine with company to ask "does have any HR openings?" instead of paging their whole list. collapse_role_group: keep one row per role_group instead of one per city, and add role_group_size saying how many postings that row stands for. Cheaper when the user wants distinct opportunities; leave it false when location or visa matters, because each city row has its own job_id and apply_channel. company: restrict to one employer. Pass whatever the user said — an id ("unitree"), or any part of the name in either language ("宇树", "Unitree", "XPeng"). Exact id/name hits win; otherwise it is a caseless substring, so a broad word can match several employers. A name this index does not carry comes back as the empty-result object with unknown_companies and suggestions. Some employers are known but deliberately NOT indexed: their careers sites run on Feishu Recruitment, whose job-list API requires a request signature; we treat that as access control and do not work around it. For those (Momenta, 小马智行 Pony.ai, 智元 AgiBot, MiniMax, 智谱 Zhipu, 商汤 SenseTime, 逐际动力 LimX, 自变量 X Square, 千寻智能 Spirit AI, 加速进化 Booster) the empty-result object carries known_not_indexed with the employer's own careers portal URL, the reason, and employer_opt_in (the employer can authorize the read-only Feishu open-platform scopes hire:site:readonly and hire:site_job_post:readonly). 蔚来 NIO is on the same list for a different reason: its own careers page publishes the postings, but its edge security policy blocks this crawler (HTTP 567) and we do not work around security controls; the employer can allowlist the crawler or authorize the same read-only scopes. Send the user to that portal; do not retry with a looser filter, and do not present the absence as "not hiring". location: caseless substring over the employer's location text, either language ("北京", "Beijing", "Mountain View", "Remote"). Alias-aware for the cities Chinese employers spell several ways: 广州 / Guangzhou also reaches rows that name only a district ("广东·天河区", 番禺区, 黄埔区, 南沙区, 海珠区, 越秀区, 白云区), 深圳 / Shenzhen reaches 南山区, 福田区, 龙岗区, 宝安区, Beijing reaches 北京市, and "remote" reaches 远程; 上海, 杭州, 苏州, 南京, 武汉, 成都 and 合肥 match their pinyin too. Combine with remote_scope to keep or drop a country. No location filter means all locations. limit: max results (default 20); must be >= 1 (else ERR_BAD_PAGE). A negative offset clamps to 0.

Salary fields are null wherever the employer's ATS publishes no range, which is most postings outside US states with pay-transparency law; require_stated_salary and currency therefore narrow mostly to US rows, and min_salary alone KEEPS unstated rows (it cannot rule them out). Pass currency to compare in one currency; without it, stated pay is compared as an annualised number in whatever currency it was stated.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
titleNo
offsetNo
remoteNo
skillsNo
companyNo
currencyNo
locationNo
min_salaryNo
role_familyNo
remote_scopeNo
required_skillsNo
collapse_role_groupNo
require_stated_salaryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the read-only/idempotent safety profile; the description adds substantial behavior it does not contain: 100-row cap returning {results, truncated, hint} instead of silent truncation, the empty-result object with results/unknown_skills/suggestions, known_not_indexed employer handling, and pagination semantics with no server-side cursor. It also explains null semantics (response_sla_days, update_signal, date_signal) rather than leaving them ambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core capability and paging/filter rules are front-loaded, but the body is heavily padded with output-interpretation prose (the long ghost_score 'how to say it' detour and multi-paragraph null discussions) that, given an output schema exists, is longer than selection guidance requires. Dense and largely on-topic, but not tightly sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter, zero-required, 0%-schema-coverage search tool with many edge cases, nothing an agent needs is missing: filter interactions, truncation, empty-result shapes, known-not-indexed employers, and paging are all covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the entire burden and does so for all 14 params: ANY-overlap vs AND semantics for skills/required_skills, alias-aware location expansion with concrete examples, title substring rules for ASCII vs CJK plus the synonym group, min_salary keeping unstated pay by default, role_family's refusal on unknown values, and limit/offset bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Search the live job index by hard filters; returns ranked JobPosting[]'), which tells the agent both the operation and the return type immediately. It also distinguishes itself from siblings by naming get_company_info as the place to get a true total for one employer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: pass company rather than filtering client-side, use collapse_role_group only when location/visa doesn't matter, use title for HR/recruiting because there is no hr family, use offset paging after collapsing, and read unknown_skills to decide between retrying with a suggestion vs loosening a filter. It even states an exclusion (do not retry with a looser filter for known_not_indexed employers).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_intentWatch a job intentA

Register a standing intent so new matches can be pulled later.

The caller supplies its OWN anonymous fingerprint (e.g. "#a3f9-k2p7-x8q1"; make it 12+ random characters, a four-character tag collides with strangers) — the client generates and owns it; the server stores but can never recover it, so persist it client-side and pass the identical one to check_watches. Only the fingerprint and non-PII filter keys are stored — never a name, email, phone or résumé. Accepted filter keys mirror search_jobs: skills (ANY-overlap), required_skills (ALL/AND — use this to keep sales / solutions-architect roles out), remote (bool), role_family (e.g. "engineering"), min_salary (int), company (one employer, resolved at registration), location (substring of the location text, alias-aware like search_jobs: 广州 also reaches 广东·天河区 rows, Beijing reaches 北京市, "remote" reaches 远程), title (list of title substrings, ANY-of, synonym-expanded for HR terms like search_jobs — the way to watch for recruiting or other roles no skill tag names). Any other key is REFUSED (ERR_UNKNOWN_FILTER) rather than silently dropped, and a role_family outside engineering | data | product | design | marketing | sales | ops | other is refused too (ERR_UNKNOWN_ROLE_FAMILY).

min_salary keeps rows with NO stated pay (they cannot be ruled out); it only drops rows whose stated pay is below the floor. Pay is stated mostly where law requires it (US postings on Greenhouse/Lever/Ashby), so a watch that needs a number will lean US.

Returns { watch_id, status, fingerprint, existing_watches, fingerprint_notice }. existing_watches > 0 means this fingerprint was already in use; if those watches are not yours, pick a longer random fingerprint.

ParametersJSON Schema
NameRequiredDescriptionDefault
filtersYes
fingerprintYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only declaring non-readOnly/non-idempotent/non-destructive, the description carries the rest and then some: privacy guarantees (server can never recover the fingerprint, no PII stored), hard refusals (ERR_UNKNOWN_FILTER for unknown keys, ERR_UNKNOWN_ROLE_FAMILY for invalid role_family), and the counterintuitive min_salary behavior of retaining rows with no stated pay.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then organized by concern (privacy, filter keys, errors, return). It is long, but the length is largely forced by 0% schema coverage and two nested free-form filters; a few parenthetical examples (e.g. the fingerprint illustration) could be trimmed without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although there is no output schema, the description enumerates the return fields (watch_id, status, fingerprint, existing_watches, fingerprint_notice) and explains how to act on existing_watches > 0. For a mutation tool with a free-form filter object and zero annotation detail, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and `filters` is a free-form object with additionalProperties, so the description must define semantics itself — and it does, documenting five accepted filter keys, their matching logic (ANY-overlap vs ALL/AND), alias-aware location matching, and enumerating legal role_family values. The fingerprint format requirement (12+ random characters) and its collision tag are also specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Register a standing intent so new matches can be pulled later') and immediately distinguishes itself from siblings by naming check_watches as the paired read tool. An agent can tell this apart from search_jobs (one-off) without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Routing is explicit: persist the fingerprint client-side and pass the identical one to check_watches; use required_skills to keep sales/solutions-architect roles out; use `title` to watch for roles no skill tag names. There is no explicit statement of when NOT to register a watch (e.g. use search_jobs for a one-time query), which keeps it just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.6.8
    • Changedsearch_jobs1 field changed
      • addedInput schema / properties / title
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Title"
        +}
  2. 1 tool updatev0.6.7
    • Changedsearch_jobs1 field changed
      • addedInput schema / properties / location
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Location"
        +}
  3. 2 tool updatesv0.6.3
    • Addedrefresh_index
    • Changedsearch_jobs2 fields changed
      • addedInput schema / properties / collapse_role_group
        Added value: +{
        +  "default": false,
        +  "title": "Collapse Role Group",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / company
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Company"
        +}
  4. 1 tool updatev0.5.0
    • Changedsearch_jobs1 field changed
      • addedInput schema / properties / offset
        Added value: +{
        +  "default": 0,
        +  "title": "Offset",
        +  "type": "integer"
        +}
  5. 5 tool updatesv0.3.2
    • First observedauthorize_application
    • First observedcheck_watches
    • First observedget_company_info
    • First observedsearch_jobs
    • First observedwatch_intent

TDQS

A4.5/5.0

Scored across 6 tools

Disambiguation4/5

The watch pair (watch_intent / check_watches) is clearly split by register-vs-pull semantics, and refresh_index, get_company_info and authorize_application each occupy a distinct role. The only mild overlap is between search_jobs filtered by `company` and get_company_info, both of which answer employer-scoped questions, but the descriptions explicitly delineate postings vs. aggregate trust signals.

Naming Consistency5/5

All six names are snake_case, verb-first constructions (watch_intent, check_watches, get_company_info, refresh_index, authorize_application, search_jobs). No camelCase drift, no vague single-word verbs, and the noun half consistently names the resource or action target.

Tool Count5/5

Six tools for a search + standing-watch + employer-trust + apply-handoff domain is well-scoped, and each one maps to a distinct stage of the workflow rather than being a near-duplicate. Nothing appears padded, and no obvious capability is crammed into an overloaded mega-tool.

Completeness4/5

Core lifecycle is covered: search, register a watch, pull new matches, force a single-employer re-crawl, inspect employer trust signals, and record an authorized application. The notable gap is watch management — there is no way to list or delete standing watches, and no direct job-detail fetch, though search results carry full posting fields so agents can work around it.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    D
    maintenance
    An MCP server that enables AI-assisted job search workflows including job discovery, application tracking, resume evaluation, and cover letter generation, with support for multiple job sources and scheduled scraping.
    83
    10 npm
    1
    AGPL 3.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    A local-first, open-source MCP server that analyzes jobs, matches your CV, tailors documents, and tracks applications — all on your machine with no data uploaded.
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that exposes job-search and application-management capabilities to compatible AI clients, enabling discovery of vacancies, drafting of tailored application materials, and coordinated human-approved submissions.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    MCP server that enables AI assistants to search LinkedIn for job posts, save them locally, and manage them via a React dashboard.
    -