Skip to main content
Glama
jhgaylor

cleanjobdata-mcp

by jhgaylor
README.md
# CleanJobData MCP Server

A Model Context Protocol (MCP) server providing tools to interact with the [CleanJobData](https://cleanjobdata.com) Job API.

[![PyPI](https://img.shields.io/pypi/v/cleanjobdata-mcp)](https://pypi.org/project/cleanjobdata-mcp/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

It runs two ways:

*   **stdio** (default) — your MCP client launches it locally and it uses the `CLEANJOBDATA_API_KEY` env var.
*   **HTTP** (`--transport http`) — one hosted server, many users, each authenticating with **their own** CleanJobData key sent per request. See [Running as a remote HTTP server](#running-as-a-remote-http-server).

## Available MCP Interactions

This server exposes the following MCP interactions:

### Tools

*   `search_jobs`: Search for jobs using the CleanJobData API based on various criteria.
    *   *Parameters*: `title`, `sort_by`, `city_id`, `state_id`, `country_id`, `location`, `remote`, `remote_type`, `company_name`, `employer_id`, `salary_min`, `salary_max`, `require_salary`, `experience_level`, `employment_type`, `published_after`, `max_age`, `include_expired`, `include_description`, `limit`, `cursor`, `count`.
*   `get_job`: Retrieve detailed information about a specific job (including its full description) by ID.
    *   *Parameters*: `job_id`.
*   `search_companies`: Search for companies by name (fuzzy), website domain, or company IDs.
    *   *Parameters*: `query`, `website_url`, `employer_id`, `active`, `limit`, `offset`.
*   `get_company`: Retrieve detailed information about a specific company, including enrichment data.
    *   *Parameters*: `company_id`.
*   `suggest_locations`: Autocomplete city/state/country names into the IDs used by `search_jobs` geo filters.
    *   *Parameters*: `query`, `kinds`, `limit`.

### Prompts

*   `create_candidate_profile`: Generates a structured prompt based on candidate details (name, LinkedIn, website, resume text) to help guide job searching.
    *   *Parameters*: `name`, `linkedin_url`, `personal_website`, `resume_text`.

## Client Setup (Examples: Claude Desktop, Cursor)

To use this server with an MCP client like Claude Desktop or Cursor, you need to configure the client to run the server process and provide the CleanJobData API key.

1.  **Ensure `uv` is installed:** `curl -LsSf https://astral.sh/uv/install.sh | sh`
2.  **Obtain a CleanJobData API Key:** Request a key from [CleanJobData](https://cleanjobdata.com). Set it as the `CLEANJOBDATA_API_KEY` environment variable.
3.  **Configure your client:**

    *   **Using `uvx`:**
        *   **Claude Desktop:** Edit your `claude_desktop_config.json`:
            ```json
            {
              "mcpServers": {
                "cleanjobdata": {
                  "command": "uvx",
                  "args": [
                    "cleanjobdata-mcp"
                  ],
                  "env": {
                    "CLEANJOBDATA_API_KEY": ""
                  }
                }
              }
            }
            ```
        *   **Cursor:** Go to Settings > MCP > Add Server:
            *   **Mac/Linux Command:** `uvx cleanjobdata-mcp`
            *   **Windows Command:** `cmd`
            *   **Windows Args:** `/c`, `uvx`, `cleanjobdata-mcp`
            *   Set the `CLEANJOBDATA_API_KEY` environment variable in the appropriate section.

    *   **Running from source (Alternative):**
        1. Clone the repo and note where you clone it to
        2. **Claude Desktop:** Edit your `claude_desktop_config.json`:
        ```json
        {
            "mcpServers": {
                "cleanjobdata": {
                    "command": "uv",
                    "args": [
                        "run",
                        "--directory",
                        "PATH_TO_REPO",
                        "cleanjobdata-mcp"
                    ],
                    "env": {
                        "CLEANJOBDATA_API_KEY": ""
                    }
                }
            }
        }
        ```

## Running as a remote HTTP server

The HTTP transport serves many users from a single process. **Each request carries its own
CleanJobData API key**, so the server holds no user credentials and every upstream call is billed
to the caller who made it.

```bash
cleanjobdata-mcp --transport http --host 0.0.0.0 --port 8000
```

The MCP endpoint is `POST /mcp` (streamable HTTP); `GET /healthz` is an unauthenticated liveness probe.

### How clients authenticate

A client sends its key on every request, in either header:

```
Authorization: Bearer <cleanjobdata-api-key>
X-CleanJobData-API-Key: <cleanjobdata-api-key>
```

`X-CleanJobData-API-Key` wins if both are present. A request with neither is rejected with a message
telling the caller how to supply one — it does **not** silently fall back to the server's own key.

Example client config (Claude Desktop / Cursor remote MCP server):

```json
{
  "mcpServers": {
    "cleanjobdata": {
      "url": "https://your-host.example.com/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_CLEANJOBDATA_API_KEY"
      }
    }
  }
}
```

Or from the command line:

```bash
npx mcp-remote https://your-host.example.com/mcp --header "Authorization: Bearer YOUR_KEY"
```

### Docker

Prebuilt images are published to GitHub Container Registry for `linux/amd64` and `linux/arm64`:

```bash
docker run --rm -p 8000:8000 ghcr.io/jhgaylor/cleanjobdata-mcp:latest
```

| Tag | Points at |
| --- | --- |
| `latest` | The most recent release |
| `0.2.0`, `0.2` | That specific release / its latest patch |
| `edge` | The tip of `main` |
| `sha-<short>` | A specific commit |

Or build it yourself:

```bash
docker build -t cleanjobdata-mcp .
docker run --rm -p 8000:8000 cleanjobdata-mcp
```

The image ships no API key — keys arrive per request. It defaults to `MCP_TRANSPORT=http`,
`HOST=0.0.0.0`, `PORT=8000`, and runs as a non-root user.

### Scaling out

Requests are stateless by default, so you can run several replicas behind a load balancer with no
sticky sessions. To use your own ASGI server with multiple workers:

```bash
uvicorn cleanjobdata_mcp.app:app --host 0.0.0.0 --port 8000 --workers 4
```

Pass `--stateful` (or `MCP_STATEFUL=1`) only if you need per-session server state; that requires
sticky routing.

### Single-tenant HTTP deployments

If you want one hosted server that always uses *your* key rather than the caller's, set both
`CLEANJOBDATA_API_KEY` and `CLEANJOBDATA_ALLOW_ENV_KEY_FALLBACK=true`. Requests that supply their own
key still use it; requests without one fall back to the server's key. Leave this off for anything
multi-user — otherwise a user who forgets their header gets billed to you.

### CLI options

| Flag | Env var | Default | Purpose |
| --- | --- | --- | --- |
| `--transport` | `MCP_TRANSPORT` | `stdio` | `stdio` or `http` |
| `--host` | `HOST` | `127.0.0.1` | Bind address (use `0.0.0.0` in a container) |
| `--port` | `PORT` | `8000` | Bind port |
| `--path` | `MCP_PATH` | `/mcp` | URL path of the MCP endpoint |
| `--json-response` | `MCP_JSON_RESPONSE` | off | Plain JSON instead of SSE, for proxies that buffer |
| `--stateful` | `MCP_STATEFUL` | off | Keep per-session state in memory |
| `--allowed-host` | `MCP_ALLOWED_HOSTS` | none | Allowed `Host` values; setting any enables DNS-rebinding protection |
| `--allowed-origin` | `MCP_ALLOWED_ORIGINS` | none | Allowed `Origin` values |

### Operational notes

*   Terminate TLS in front of the server (load balancer, reverse proxy, or platform ingress). Keys
    travel in request headers, so plain HTTP over the public internet would expose them.
*   Tool handlers run in a thread pool, so a slow upstream call blocks one thread rather than the
    whole event loop. Very high concurrency benefits from more replicas rather than one large process.

## Development

This project uses:
- `uv` for dependency management and virtual environments
- `ruff` for linting and formatting
- `hatch` as the build backend

### Common Tasks

```bash
# Setup virtual env
uv venv

# Install dependencies
uv pip install -e .

# install cli tools
uv tool install ruff

# Run linting
ruff check .

# Format code
ruff format .
```

## Environment Variables

-   `CLEANJOBDATA_API_KEY`: Your API key for the CleanJobData API, sent upstream as a Bearer token.
    Required for stdio; over HTTP the caller's own header supplies the key instead.
-   `CLEANJOBDATA_ALLOW_ENV_KEY_FALLBACK`: Set to `true` to let HTTP requests without a key fall back
    to `CLEANJOBDATA_API_KEY`. Off by default — see
    [Single-tenant HTTP deployments](#single-tenant-http-deployments).
-   `CLEANJOBDATA_API_BASE`: Override the API base URL (default `https://api.cleanjobdata.com`).

Transport settings (`MCP_TRANSPORT`, `HOST`, `PORT`, …) are listed under [CLI options](#cli-options).

## Testing

This project uses `pytest` for testing the core tool logic. Tests mock external API calls using `unittest.mock`.

1. Install test dependencies:
```bash
# Ensure you are in your activated virtual environment (.venv)
uv pip install -e '.[test]'
```

2. Run tests:
```bash
pytest
```

## Contributing

Contributions are welcome.

## License

This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.

TDQS

A4.4/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a distinct purpose: searching jobs, retrieving job details, retrieving company details, searching companies, and suggesting location filters. No two tools overlap in functionality, making selection unambiguous.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case: search_jobs, get_job, get_company, search_companies, suggest_locations. The naming is predictable and clearly indicates the action and resource.

Tool Count5/5

Five tools is well-scoped for a job search API. Each tool covers a core operation (search, detail retrieval, and location lookup) without redundancy or unnecessary bloat.

Completeness5/5

For a read-only job search service, the tool surface covers all necessary workflows: finding jobs, inspecting job details, finding companies, inspecting company details, and resolving location IDs for filtering. There are no obvious gaps that would hinder an agent.

Maintenance

ActivitySlowing
ResponsivenessNo issues