Skip to main content
Glama
eliottreich

Crawdar Business Research

README.md
# Crawdar MCP

Crawdar gives AI agents evidence-backed public business research through a hosted MCP server. It separates verified fits, likely fits, and excluded candidates, keeps source URLs attached, and explains why candidates were rejected.

No package or local crawler is required.

## Optional local stdio bridge

This repository also provides an MIT-licensed, runnable bridge for clients that require stdio. It uses pinned `mcp-remote` to forward MCP traffic to `https://crawdar.com/api/mcp`. The crawler backend remains hosted and is not included in this repository. This is not an offline or self-hosted crawler.

With Node.js 22 or newer, run `npm ci --ignore-scripts`, then configure your MCP client to launch `node` with the absolute path to `node_modules/mcp-remote/dist/proxy.js`, followed by `https://crawdar.com/api/mcp --transport http-only` as separate arguments. For a container, build with `docker build -t crawdar-mcp .` and launch with `docker run --rm -i crawdar-mcp` (no TTY).

The bridge exposes the hosted server's tools and sends tool inputs to Crawdar over HTTPS. Live research remains subject to Crawdar's quotas and service availability. Credentials are not embedded. The fictional sandbox and tool discovery require no key. For authenticated usage, use the direct HTTP configuration below or the bridge's `--header-file` option with a private local header file. Never commit credentials or include them in command arguments.

## Bring a difficult search brief

We are recruiting five independent builders to test a real company-research workflow. [Describe your brief and acceptance rules](https://github.com/eliottreich/crawdar-mcp/issues/new?template=workflow-test.yml). We will confirm a small sample scope first, then review what matched, what failed and what remains unknown. No stars, referrals or positive reviews required.

GitHub issues are public. Keep confidential briefs and all credentials out of them; use hello@crawdar.com for private support. The general [Crawdar web tool](https://crawdar.com/) does not require a GitHub account.

## Connect

Remote MCP endpoint:

```text
https://crawdar.com/api/mcp
```

Claude Code:

```bash
claude mcp add --transport http crawdar https://crawdar.com/api/mcp
```

Codex configuration:

```toml
[mcp_servers.crawdar]
url = "https://crawdar.com/api/mcp"
```

Generic MCP configuration:

```json
{
  "mcpServers": {
    "crawdar": {
      "type": "http",
      "url": "https://crawdar.com/api/mcp"
    }
  }
}
```

## Tools

- `research_businesses`: One-call research from a natural-language brief.
- `search_businesses`: Structured target and geography search.
- `sandbox_businesses`: Fictional integration test with no search-provider usage.
- `start_lead_search`: Start a private resumable job.
- `get_lead_search`: Poll status and page through compact or full results.
- `refine_lead_search`: Create a revised job without changing the original.
- `retry_lead_search`: Recover from a stopped job.
- `cancel_lead_search`: Cancel a queued or running job.
- `explain_crawdar`: Read result semantics, limits, and safe-use guidance.

## Try a visual workflow in n8n

Download [the n8n sandbox workflow](examples/n8n-sandbox.json), import it into your n8n instance, and click **Execute Workflow**. This demonstrates routing, not real prospect discovery. No Crawdar account, key, AI-model subscription, or CRM credentials are needed. You still need n8n itself.

It fetches three fictional businesses, keeps source evidence and sandbox labels attached, and sends two explicitly qualified examples down one branch and one uncertain example to review. No emails, CRM writes or scheduled scans. Changing the brief only changes fixture labels, not real research. A qualified status is not a guarantee of fit: review the underlying evidence before consequential use.

Verified September 6, 2026 by importing and executing the exact JSON in n8n 2.37.10 with Node 24.20.0. Assertions passed for the 2/1 branch split, retained evidence and sandbox markers, and zero search-provider requests. This is a founder-maintained example, not an n8n-approved gallery template or a live-search accuracy benchmark.

For real research, follow the [agent guide](https://crawdar.com/agents) and use authenticated durable searches with rate-limit handling. Do not simply change the sandbox URL and assume the workflow handles job polling, pagination or paid usage.

## Free agent identity

Before creating a key, you can inspect the API with the [Postman sandbox collection](examples/crawdar-sandbox.postman_collection.json). Import the JSON into Postman, or run it locally with Node.js and [Newman](https://learning.postman.com/docs/reference/newman-cli/command-line-integration-with-newman):

```bash
npx --yes newman@6.2.1 run https://raw.githubusercontent.com/eliottreich/crawdar-mcp/main/examples/crawdar-sandbox.postman_collection.json --timeout-request 20000 --timeout 60000 --bail
```

The command downloads the pinned Newman runner and public collection. No Crawdar or Postman account is needed for this CLI route. The collection only calls the fictional sandbox, makes three requests, and checks the response format, preserved uncertainty, source links and validation errors. It does not measure real lead quality or send messages. Verified September 6, 2026 with Newman 6.2.1: 13 assertions passed. Postman desktop import has not been separately tested.

Create an optional free API key:

```bash
curl -X POST https://crawdar.com/api/v1/keys \
  -H 'content-type: application/json' \
  -d '{"label":"My research agent"}'
```

Pass the returned key as `x-api-key` or `Authorization: Bearer`. The secret is shown once.

## Contracts and documentation

- [Portable research skill](skills/crawdar-business-research/SKILL.md): instructions for evidence review, private tokens, error handling and exports. No automatic outreach.
- [Smithery listing](https://smithery.ai/servers/eliottreich/crawdar): optional distribution route. Smithery requires its own authentication; a Crawdar key is separate. Its Connect API was tested with fictional sandbox results on September 6, 2026. CLI 1.2.0 connection commands returned 404 in that check, so use the direct MCP endpoint above or Smithery's documented Connect API.

- [Agent guide](https://crawdar.com/agents)
- [OpenAPI 3.1](https://crawdar.com/openapi.json)
- [llms.txt](https://crawdar.com/llms.txt)
- [Portable agent skill](https://crawdar.com/skills/crawdar-business-research/SKILL.md)
- [Crawler and responsible-use policy](https://crawdar.com/bot)
- [Official MCP Registry listing](https://registry.modelcontextprotocol.io/v0.1/servers?search=com.crawdar%2Fbusiness-research)

Public information can be incomplete or outdated. Keep evidence URLs attached and verify records before outreach or another consequential action.

TDQS

A4.3/5.0

Scored across 9 tools

Disambiguation4/5

Most tools have clear, distinct purposes: synchronous vs asynchronous search, sandbox, and explain. The overlap between research_businesses and search_businesses is somewhat subtle but resolved by the explicit distinction between natural-language briefs and structured criteria.

Naming Consistency5/5

All tool names follow a consistent snake_case action_object pattern (research_businesses, start_lead_search, cancel_lead_search), making their behavior predictable. Even explain_crawdar fits the convention.

Tool Count5/5

Nine tools is well-scoped for a business-research service: it covers synchronous search, async job lifecycle, sandbox testing, and explanatory metadata without excess or redundancy.

Completeness5/5

The tool surface covers the full search workflow—immediate sync queries, durable async jobs, refinement, retry, cancellation, status/result reading, and safe testing—plus an explain tool for interface guidance. No obvious gaps remain.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive