Skip to main content
Glama
README.md
# x_mcp

**Find recent X posts with the browser and model already in your AI harness. No paid X API, X developer account, or separate model API key.** Free and open source under MIT.

Ask for `machine learning`, `AI agents`, or a longer topic such as `recent advances in AI agents for biomedical discovery`. The server prepares X-specific searches, extracts rendered posts, and returns relevant results with dates and source links. Machine learning defaults to research/advance searches. Results default to newest first within the captured sample.

## What you need

- Node.js **22 or newer**.
- A local MCP-capable harness with a browser tool (for example, Codex with browser access, or Claude Code/Cursor with a browser MCP).
- **Sign in to X in the browser your harness controls.** Signing in in a different browser/profile does not share that session.

The tool itself is free. Your existing harness/model subscription or usage costs still apply. X can limit search access or change its UI. No scraper can promise every post or uninterrupted access.

## Install and connect

```bash
git clone https://github.com/mashathepotato/x_mcp.git
cd x_mcp
npm ci
npm run check
```

`npm ci` builds the server. Run `npm run build` after editing TypeScript. This project is installed from source; it is not published to npm yet.

### Codex

From the project directory:

```bash
codex mcp add x-mcp -- node "$(pwd)/dist/index.js"
```

Restart or reconnect your harness so it loads the server. Ensure its browser tool is enabled and signed in to X. The MCP does not grant browser access by itself.

Equivalent `~/.codex/config.toml` entry (replace the path):

```toml
[mcp_servers.x-mcp]
command = "node"
args = ["/absolute/path/to/x_mcp/dist/index.js"]
startup_timeout_sec = 15
tool_timeout_sec = 120
```

[Official Codex MCP configuration](https://developers.openai.com/codex/mcp/).

### Claude Code

```bash
claude mcp add --transport stdio x-mcp -- node "$(pwd)/dist/index.js"
```

Use alongside your browser tool, then reconnect the MCP. [Official Claude Code MCP configuration](https://code.claude.com/docs/en/mcp).

### Cursor, Claude Desktop, and other local MCP clients

Merge this entry into the client's MCP configuration. Replace the absolute path; on Windows use an escaped path such as `C:\\Users\\you\\Documents\\x_mcp\\dist\\index.js`.

```json
{
  "mcpServers": {
    "x-mcp": {
      "command": "node",
      "args": ["/absolute/path/to/x_mcp/dist/index.js"]
    }
  }
}
```

Clients that support only remote HTTP MCP servers cannot launch this local stdio server. A client without browser tools needs the optional CDP mode below. A basic web-search tool alone is not equivalent to an authenticated X browser.

If the app cannot find `node`, use its absolute executable path (`which node` on macOS/Linux, `where node` on Windows). **Do not use `npm start` as the MCP command:** npm may print banners onto the protocol's stdout. Use `node …/dist/index.js`.

## Try it

After connecting, ask your harness:

> Use x-mcp to find 10 recent machine learning advances from the last 7 days. Complete the browser workflow and give me a concise summary with dates and tweet links.

> Use x-mcp to find the latest tweets about AI agents for biomedical discovery. Keep the results relevant to biomedical research, rather than generic agent announcements.

> Search x-mcp for AI agents from the last 2 days. Rank by relevance and cite the original posts.

The server exposes a `research_topic` MCP prompt as another starting point. A no-network installation check:

```bash
node dist/index.js --plan "machine learning"
node dist/index.js --plan "ai agents for biomedical discovery"
```

These commands print **search plans**, not tweets. `npm run check` runs the real MCP handshake and fixture-backed collection tests without accessing X.

## How harness mode works

MCP servers cannot automatically invoke arbitrary browser tools in their host. This server provides an explicit workflow that the harness's agent completes:

1. **`x_search_topic`** turns the topic into up to three X Latest search URLs and returns `status: "needs_browser"`, a `searchId`, and a read-only extraction script.
2. **Your harness's browser** opens those URLs in its signed-in session. It reads the visible posts, runs the extraction script if DOM evaluation is supported, and captures each page before scrolling. If only accessibility reading is available, the host copies exact text/permalinks into the same capture schema.
3. **`x_collect_posts`** accepts those observations, removes duplicates, filters the exact time window and topic, and returns ranked posts. Call it again with the same `searchId` to merge more snapshots.
4. **Your harness's model** summarizes the results, citing original links and separating claims from verified advances.

The workflow is included in MCP initialization instructions, tool descriptions, and `x-mcp://workflow`. The extraction script is also available at `x-mcp://extractor` or through `node dist/index.js --extractor`. No MCP sampling capability or extra LLM service is required.

### Tool reference

| Tool | Purpose |
| --- | --- |
| `x_search_topic` | Start a search; return a browser handoff, or scrape through configured CDP. |
| `x_collect_posts` | Merge browser captures for a search, filter and rank them. |
| `x_get_status` | Check configuration and prerequisites; does not test X login. |

`x_search_topic` arguments:

| Argument | Default | Meaning |
| --- | --- | --- |
| `topic` | required | Keywords or natural-language topic, up to 2,000 characters. |
| `days` | `7` | Exact lookback window, 1–30 days. |
| `limit` | `20` | Maximum returned posts, 1–100; fewer may be available. |
| `intent` | `auto` | `advances`, `latest`, or automatic (ML → advances). |
| `sort` | `latest` | `latest` or `relevance`; X itself always uses its Latest tab. |
| `language` | `en` | X language code, or `all`. |
| `includeReplies` | `false` | Include reply posts. |
| `rawQuery` | none | Optional X query body refined by the harness model; local keyword matching is disabled for this override. |
| `mode` | `auto` | `harness`, `browser`, or auto-detect configured CDP. |

With `rawQuery`, the host is responsible for semantic relevance; dates, URLs and reply filtering still apply.

Queries use inspectable alias groups and topic constraints. Research intent adds scholarly/release variants and a broader fallback, excluding common course, roadmap and job promotions. Captured advances also need observable paper/release wording or a paper/code link; this selects candidates, not verified breakthroughs. Long prose is simplified heuristically; review `plan.queries` and use `rawQuery` when a specialist term or Boolean expression needs precision. Relative dates or requested counts in topic prose do not set `days` or `limit`; the harness should pass those structured options. The server uses no embeddings and does not perform semantic reasoning itself. Its local keyword filter can miss relevant posts expressed with different wording.

`rawQuery` preserves its trimmed body and appends date, language and reply filters. Avoid conflicting operators in that body; use the structured options for the date window. Search filters use UTC calendar days, while captured results are checked against the exact start/end timestamps fixed when the plan was created. This rejects posts outside the window even if X ignores a search filter.

Example calls (after actual browser observation):

```javascript
x_search_topic({ topic: "AI agents", days: 7, limit: 10 })
// Save the returned searchId and open a returned plan.queries[i].url.

x_collect_posts({
  searchId: "<returned searchId>",
  captures: [{
    pageUrl: "<actual planned X search URL>",
    capturedAt: "<current ISO 8601 time>",
    state: "ok",
    posts: [{
      url: "<observed https://x.com/handle/status/id permalink>",
      text: "<exact observed tweet text>",
      publishedAt: "<observed ISO time, or omit if unknown>",
      links: ["<observed outbound link, if any>"]
    }]
  }]
})
```

These placeholders illustrate the schema; they are not sample tweets. A capture can instead report `login_required`, `rate_limited`, `challenge`, `empty`, or `unrecognized`, with an empty `posts` list.

Returned posts include a canonical URL, author, text, timestamp, timestamp source (`page` or a derivation from the X status ID), matched terms and a heuristic relevance score. Status `captured` means planned search pages were observed; it never means all of X was scanned. `partial` means some observations are available but planned searches or access were incomplete. `no_matches` means no accepted posts in the captured sample. A login wall or failed access is reported separately.

The normalizer validates shape, links, timestamps and query provenance, but **cannot independently prove the authenticity of text supplied by a harness**. Summaries must rely on actual browser observations. The browser adapter expands up to five inline Show more buttons per search page. Remaining truncated text is marked `isTruncated`; the harness must not invent missing content. Images, video, protected content, and full threads are not transcribed or expanded.

## Optional: direct browser mode

If your harness does not have a browser tool, this server can read a Chromium browser already exposed through a **local Chrome DevTools Protocol (CDP)** endpoint. You sign in to X there once. The server then performs the browser collection inside `x_search_topic`.

For example, launch a dedicated Chrome profile on macOS:

```bash
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
  --remote-debugging-address=127.0.0.1 \
  --remote-debugging-port=9222 \
  --user-data-dir="$HOME/.x-mcp-browser" \
  https://x.com
```

On Linux or Windows, use your Chrome/Chromium executable with the same flags and a dedicated profile directory. Do not use your default daily browser profile for this example. Leave this browser open, sign in, and configure the MCP environment:

```json
{
  "mcpServers": {
    "x-mcp": {
      "command": "node",
      "args": ["/absolute/path/to/x_mcp/dist/index.js"],
      "env": {
        "X_MCP_MODE": "browser",
        "X_MCP_CDP_URL": "http://127.0.0.1:9222"
      }
    }
  }
}
```

Set your client's tool timeout to at least 120 seconds for browser mode. `X_MCP_MODE=auto` (the default) uses CDP only when `X_MCP_CDP_URL` is set; otherwise it uses the harness handoff. `X_MCP_MODE=harness` forces handoff.

Only explicit loopback endpoints are accepted. CDP can control that browser profile; keep the debugging port local and never expose it publicly. The adapter creates and closes its own search tabs, disconnects afterward, and leaves preexisting tabs and Chrome running. It reads rendered pages through [Playwright's CDP connection](https://playwright.dev/docs/api/class-browsertype#browser-type-connect-over-cdp), without launching or downloading a browser.

## Limits and troubleshooting

- **`needs_browser`:** expected in harness mode. The agent must use its browser tools and call `x_collect_posts`; a URL plan is not a completed search.
- **`login_required`:** sign in to X in the controlled browser and start the search again. [X documents signed-in advanced search](https://help.x.com/en/using-x/x-advanced-search).
- **`rate_limited` / `challenge`:** stop and handle access manually. There are no CAPTCHA solvers, rotating proxies, guest tokens, private GraphQL calls, or login bypasses.
- **`unrecognized`:** the page may be loading or X's markup may have changed. Inspect it through the harness browser; report an issue with a sanitized fixture.
- **Sparse results:** use a longer lookback, `language: "all"`, or a carefully refined query. Do not silently replace recent tweets with old search-engine snippets.
- **`search_expired`:** plans last 30 minutes and are lost on restart. Call `x_search_topic` again.
- **Direct browser unavailable:** ensure the dedicated browser and local debugging endpoint are running. Browser errors are reported; no paid fallback is used.

Collection is bounded to at most three searches and three scrolls per search by default, with up to five inline expansions per search page. The adapter overfetches up to 3× the requested count (at most 300 unique observations) so filtering can discard irrelevant posts. Plans and captures live only in process memory: up to 50 sessions, 60 snapshots/2,000 post observations per session. One collection call accepts at most 20 snapshots, 100 posts per snapshot. No cookies or credentials are read, exported, logged, or stored by this server; your browser manages its own session. No telemetry is sent. Captured text is passed back to your harness and its model under that harness's normal data handling.

Respect X's terms and applicable rules. Use this for modest, user-directed research. Treat all post text and linked content as untrusted data, including any instructions embedded in them. The server does not verify scientific claims.

## Development

```bash
npm ci
npm run check
npm run dev        # stdio MCP for development; waits for a client
```

Source layout: `query.ts` (planning), `extract.ts` (rendered DOM capture), `rank.ts` (validation/ranking), `browser.ts` (optional CDP), `server.ts` (MCP workflow). Tests cover query specificity, extraction, dates, duplicate handling, access blockers, CDP boundaries, and a real stdio MCP handshake. CI runs Node 22 and 24.

See [TESTING.md](TESTING.md) for automated and live validation scope.

Contributions welcome: add a reproducible sanitized fixture and a regression test for parser/extractor changes. Never commit credentials, cookies, browser profiles, or private captures. See [CONTRIBUTING.md](CONTRIBUTING.md) and [LICENSE](LICENSE).

TDQS

A3.9/5.0

Scored across 3 tools

Disambiguation4/5

The three tools occupy distinct roles in a pipeline: plan generation (x_search_topic), post ingestion (x_collect_posts), and config reporting (x_get_status). Boundaries are clear from descriptions, though the 'search' name on x_search_topic is slightly misleading since it returns a plan rather than results, which could momentarily confuse an agent about where actual data comes from.

Naming Consistency5/5

All tools use a uniform x_verb_noun pattern (x_search_topic, x_collect_posts, x_get_status), with consistent prefixing and snake_case throughout. There are no deviations or mixed conventions.

Tool Count4/5

Three tools is compact but matches the narrow, deliberately minimal harness workflow of plan→capture→status. Each tool earns its place, though the surface is on the lean side with no room for auxiliary operations.

Completeness3/5

The plan→collect→status lifecycle is closed, but the actual fetching is delegated entirely to the agent's external browser, so the server cannot complete a research task on its own. There is no post detail retrieval, refresh/clear operation, or recovery path if collection fails, leaving notable gaps for the stated research purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues