Skip to main content
Glama
gzchenhao

OpenHire — Real Job Postings, Ghost Jobs Scored

Search jobs

search_jobs
Read-onlyIdempotent

Filter live job postings by skills, company, location, salary, and remote scope to get ranked results with ghost scores and employer ATS dates.

Instructions

Search the live job index by hard filters; returns ranked JobPosting[].

The server does ONLY a hard filter plus a fixed ranking of match-quality × freshness — precise re-ranking is left to you, the client, which holds the user's context. Every result includes the five protocol fields (verified_at, source, ghost_score, response_sla_days, apply_channel) plus datePosted, days_open, remote_scope, eligible_regions, role_group and ghost_reason.

verified_at and ghost_score answer DIFFERENT questions and routinely disagree: a posting confirmed live today can score 1.0. Live means the employer's ATS still returns it; the score means it has been returned for a long time, or keeps being relisted. ghost_reason says which input drove the score ("age only: open 367d, never relisted") so you can tell the user that instead of a bare number. Neither field measures intent — a long-open role can equally mean hard-to-fill.

How to say it: ghost_score is a measurement of time on the market, never a verdict that a posting is fake. Do not call a posting a ghost, zombie, fake job or 僵尸岗 on the strength of it, and do not tell the user "don't apply" because of it. Report what was measured — "open 1,616 days, never relisted, the ATS reports no last-touched date" — and let the user weigh it. A 1.0 can be an abandoned req or a role that has been genuinely hard to fill for four years; the number cannot tell those apart.

response_sla_days is null on almost every row, and that is a meaningful null: it is the employer's OWN committed reply window, set only when they claim their tenant. We never infer or estimate it — the application deep-links to the employer and never touches this server, so we cannot observe a reply even in principle. Read null as "no employer has claimed this tenant", never as "missing data" or "slow".

employer_correction appears only when the employer has claimed this tenant and said something about this specific role. It is the employer's own account, verified by corporate identity, and it sits BESIDE ghost_score, never on it: status: "evergreen" explains a high score (they hire continuously, so there is no single opening to fill), it does not lower it. status: "closed" means they say they are no longer hiring even though their ATS still returns the row, so stop sending people there. date_semantics: "requisition_created" means their ATS reports the day the req was opened internally rather than the day it went live, so days_open overstates for that employer. Treat all of it as the employer's claim, attributed, not as our measurement.

days_since_update is the third date and the one that usually settles it: the employer's own last-touched timestamp from their ATS. Two rows can both score 1.0 and mean opposite things — open 367d and untouched for 367d reads as abandoned; open 327d but touched 13 days ago reads as a tended evergreen req. Among rows where the ATS reports one at all, the share of ghost>=0.99 postings touched by the employer inside 30 days has ranged from about a third to three quarters across refreshes; read the current value from pct_ghost_hi_touched_within_30d in docs/numbers.json rather than quoting a figure.

Null is NOT "abandoned": Ashby, Lever and Beisen do not report a last-touched date, and those rows carry update_signal: "not_reported_by_ats" instead. A row carrying date_signal: "not_reported_by_ats" goes one step further: its source reports no posting date at all (Li Auto's first-party mirror is one), so its datePosted and days_open are counted from the day this index first saw it, not from the employer's own date. Read those as a lower bound on age, never as the employer's timeline. Treating that null as "nobody has touched this in a year" would describe the vendor's API, not the employer. Honest limit even when present: an ATS bumps it on any edit or re-publish, so it means "touched", not necessarily "content changed".

To answer "what is hiring?", pass company — do not filter client-side.

Skills are alias-aware on the request side: 感知 finds rows tagged perception, 占用网络 finds occ / occupancy / occupancy networks. If a requested tag matches NOTHING in the index, the response switches to {results, unknown_skills, suggestions} so you can tell the user which word was dropped instead of presenting a half-match as a match.

role_family filters exclude rows KNOWN to be another family; rows not yet classified (role_family null, typically the newest postings) are included, not hidden. remote_scope can also be "unknown": remote with a location we could not read; it is never reported as "worldwide" any more.

Two things worth knowing before you spend your budget:

  • role_group is shared by the same role posted in several cities — one employer may list one job 22 times, once per location. Those rows are genuinely distinct (each has its own job_id and apply_channel, which matters when the user has a location or visa constraint), but if you only need distinct opportunities, group by role_group and keep one per group. Measured, about 20% of a page is same-role repeats.

  • offset pages through the ranked list. After collapsing by role_group, call again with offset += limit to get more distinct roles. Fewer rows than limit means you reached the end. There is no server-side cursor to keep alive.

  • One call returns at most 100 rows. Ask for more and you get an object with results and truncated, not a silent first page: 100 of 198 looked exactly like "this employer has 100 jobs". truncated: true means a full page came back and there may be more, with the offset to call next in hint; truncated: false means fewer rows than the page size matched and there is no next page. For one employer, get_company_info's active_jobs is the true total.

  • Skill tags match separator-insensitively: "computer vision", "computer-vision" and "computer_vision" are one skill. The extractor emits all three spellings, so an exact-string search reached as little as 42% of the rows that had the skill.

  • An empty search does NOT return a bare list. It returns an object with results: [] plus hint, unknown_skills and suggestions, because [] alone cannot tell you whether you mistyped a tag or the market is genuinely dry. Read unknown_skills: if it is non-empty those tags exist nowhere in the index and you should retry with a suggestion; if it is empty your tags were fine and you should loosen a filter.

Args: skills: skill tags, ANY-overlap match (union), e.g. ["rust", "k8s"]. required_skills: skills that must ALL be present (AND), e.g. ["rust"]. remote: if true, only fully-remote roles. remote_scope: filter remote roles by reach: "worldwide" | "unknown" | "region_locked" | "country_locked". min_salary: salary floor. By default roles with NO stated pay are KEPT (they can't be ruled out); set require_stated_salary=true to drop them. currency: restrict to a stated-pay currency, e.g. "USD" (implies stated pay). require_stated_salary: if true, drop roles that publish no salary. role_family: coarse family filter, one of engineering | data | product | design | marketing | sales | ops | other. Any other value is refused (ERR_UNKNOWN_ROLE_FAMILY) rather than ignored: "recruiting" used to pass through and return the whole index. There is NO hr / recruiting / people family — those roles are filed under ops; use title to isolate them. Populated for most live rows, so this is an effective way to keep sales / solutions-architect roles out of an engineering search. title: caseless substrings over the job title, ANY-of, e.g. ["recruit", "招聘"] or ["感知"]. This is the filter for roles no skill tag or family can isolate: HR, recruiting, finance, legal, a specific team name. An ASCII term must start a word ("hr" reaches HR, HRBP and HR Business Partner but not Chrome); a CJK term is a plain substring. One synonym group is expanded for you: any of hr / hrbp / human resources / recruit / talent acquisition / sourcer / people ops / 人力 / 人事 / 招聘 reaches all the others, so title=["招聘"] also finds an English "Senior Technical Recruiter". Combine with company to ask "does have any HR openings?" instead of paging their whole list. collapse_role_group: keep one row per role_group instead of one per city, and add role_group_size saying how many postings that row stands for. Cheaper when the user wants distinct opportunities; leave it false when location or visa matters, because each city row has its own job_id and apply_channel. company: restrict to one employer. Pass whatever the user said — an id ("unitree"), or any part of the name in either language ("宇树", "Unitree", "XPeng"). Exact id/name hits win; otherwise it is a caseless substring, so a broad word can match several employers. A name this index does not carry comes back as the empty-result object with unknown_companies and suggestions. Some employers are known but deliberately NOT indexed: their careers sites run on Feishu Recruitment, whose job-list API requires a request signature; we treat that as access control and do not work around it. For those (Momenta, 小马智行 Pony.ai, 智元 AgiBot, MiniMax, 智谱 Zhipu, 商汤 SenseTime, 逐际动力 LimX, 自变量 X Square, 千寻智能 Spirit AI, 加速进化 Booster) the empty-result object carries known_not_indexed with the employer's own careers portal URL, the reason, and employer_opt_in (the employer can authorize the read-only Feishu open-platform scopes hire:site:readonly and hire:site_job_post:readonly). 蔚来 NIO is on the same list for a different reason: its own careers page publishes the postings, but its edge security policy blocks this crawler (HTTP 567) and we do not work around security controls; the employer can allowlist the crawler or authorize the same read-only scopes. Send the user to that portal; do not retry with a looser filter, and do not present the absence as "not hiring". location: caseless substring over the employer's location text, either language ("北京", "Beijing", "Mountain View", "Remote"). Alias-aware for the cities Chinese employers spell several ways: 广州 / Guangzhou also reaches rows that name only a district ("广东·天河区", 番禺区, 黄埔区, 南沙区, 海珠区, 越秀区, 白云区), 深圳 / Shenzhen reaches 南山区, 福田区, 龙岗区, 宝安区, Beijing reaches 北京市, and "remote" reaches 远程; 上海, 杭州, 苏州, 南京, 武汉, 成都 and 合肥 match their pinyin too. Combine with remote_scope to keep or drop a country. No location filter means all locations. limit: max results (default 20); must be >= 1 (else ERR_BAD_PAGE). A negative offset clamps to 0.

Salary fields are null wherever the employer's ATS publishes no range, which is most postings outside US states with pay-transparency law; require_stated_salary and currency therefore narrow mostly to US rows, and min_salary alone KEEPS unstated rows (it cannot rule them out). Pass currency to compare in one currency; without it, stated pay is compared as an annualised number in whatever currency it was stated.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNo
titleNo
offsetNo
remoteNo
skillsNo
companyNo
currencyNo
locationNo
min_salaryNo
role_familyNo
remote_scopeNo
required_skillsNo
collapse_role_groupNo
require_stated_salaryNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.6.8
    • addedInput schema / properties / title
      Added value: +{
      +  "anyOf": [
      +    {
      +      "items": {
      +        "type": "string"
      +      },
      +      "type": "array"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Title"
      +}
  2. Changed1 schema field changedv0.6.7
    • addedInput schema / properties / location
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Location"
      +}
  3. Changed2 schema fields changedv0.6.3
    • addedInput schema / properties / collapse_role_group
      Added value: +{
      +  "default": false,
      +  "title": "Collapse Role Group",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / company
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Company"
      +}
  4. Changed1 schema field changedv0.5.0
    • addedInput schema / properties / offset
      Added value: +{
      +  "default": 0,
      +  "title": "Offset",
      +  "type": "integer"
      +}
  5. First observedv0.3.2

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the read-only/idempotent safety profile; the description adds substantial behavior it does not contain: 100-row cap returning {results, truncated, hint} instead of silent truncation, the empty-result object with results/unknown_skills/suggestions, known_not_indexed employer handling, and pagination semantics with no server-side cursor. It also explains null semantics (response_sla_days, update_signal, date_signal) rather than leaving them ambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core capability and paging/filter rules are front-loaded, but the body is heavily padded with output-interpretation prose (the long ghost_score 'how to say it' detour and multi-paragraph null discussions) that, given an output schema exists, is longer than selection guidance requires. Dense and largely on-topic, but not tightly sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter, zero-required, 0%-schema-coverage search tool with many edge cases, nothing an agent needs is missing: filter interactions, truncation, empty-result shapes, known-not-indexed employers, and paging are all covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the entire burden and does so for all 14 params: ANY-overlap vs AND semantics for skills/required_skills, alias-aware location expansion with concrete examples, title substring rules for ASCII vs CJK plus the synonym group, min_salary keeping unstated pay by default, role_family's refusal on unknown values, and limit/offset bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Search the live job index by hard filters; returns ranked JobPosting[]'), which tells the agent both the operation and the return type immediately. It also distinguishes itself from siblings by naming get_company_info as the place to get a true total for one employer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: pass company rather than filtering client-side, use collapse_role_group only when location/visa doesn't matter, use title for HR/recruiting because there is no hr family, use offset paging after collapsing, and read unknown_skills to decide between retrying with a suggestion vs loosening a filter. It even states an exclusion (do not retry with a looser filter for known_not_indexed employers).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.