chat-mcp
This server delegates code reviews, design comparisons, and debugging to your signed-in ChatGPT web session via Chrome, using local files for context.
chatgpt_ask— Send a prompt to ChatGPT, optionally attaching local files/folders (context_paths) or a repo root (repo_path) for context collection; supports follow-ups viaconversation_handle.review_with_chatgpt— Run substantial code reviews by collecting the diff and file contents directly fromrepo_pathwithscopeofworking_tree,staged, orbranch(plus optionalpaths,base_ref,question); returns actionable findings.chatgpt_health— Check readiness before the first request; returnsreadyplus anext_actionfor setup, connection, or recovery without exposing chat content. (Read-only, no side effects.)chatgpt_cancel— Cancel an in-flight generation or clear an unchanged unsent draft for a givenrequest_id.Fault-tolerant retries — Every call takes a
request_id; retry uncertain outcomes with the same ID and identical inputs, or retrieve results later without resending the prompt.No pasting needed — Context is gathered locally from your repo, so large codebases can be reviewed without copying code into the conversation.
Allows Codex to perform code reviews and analysis through ChatGPT by automating a connected Chrome tab to send selected code/diffs and retrieve answers.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@chat-mcpreview my changes"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
chat-mcp
A local Codex plugin for code reviews, design comparisons, and debugging through your ChatGPT account in Chrome.
Review and analysis requests run in your signed-in ChatGPT web tab. The plugin collects selected files locally and returns the answer to Codex, so the delegated analysis uses ChatGPT web access while keeping that source context out of the coordinating Codex conversation. Browser UI changes can still require compatibility updates; the recovery state below makes those failures inspectable without duplicate sends.
Install
Requires Codex, Node.js 22+ (with npm), Git, Chrome 116+, and a ChatGPT account.
git clone https://github.com/phykn/chat-mcp.git
cd chat-mcp
npm run setupSetup installs dependencies, builds and registers the plugin, prepares the Chrome extension, and checks the connection. It prepares the Codex CLI if needed. No separate marketplace registration is required.
Connect Chrome once
Setup opens Chrome's extension settings and the prepared extension folder in an interactive terminal. If needed, open chrome://extensions manually.
Enable Developer mode and click Load unpacked.
Select the exact folder printed by setup. Paste the path into the folder picker’s address bar on Windows, or use Cmd+Shift+G on macOS.
Sign in to the ChatGPT tab that opens. Select Chat if prompted.
Return to the terminal and press Enter to verify the connection.
The default folder is .chat-mcp/extension under your home directory. Do not select the repository's extension folder or a manifest.json file. The prepared folder contains the generated configuration.
To finish later, enter q, then run npm run doctor when ready.
Related MCP server: chatgpt-mcp
Use
Open a new Codex task after installation and ask:
Check my Chat MCP connection.
Then:
Use Chat MCP to review my changes.
You can also ask for a structure or readability review of existing code. review_with_chatgpt with scope: working_tree accepts file or folder paths and includes their current contents even without a diff. Without paths, it prefers changed files, or reviews current tracked files when there are no eligible changes. staged and branch remain limited to changes in those scopes.
Keep the connected ChatGPT tab open and your PC awake during requests. Requests activate that tab and restore its Chrome window if minimized. The bridge starts when needed; Chrome reconnects automatically after restarting. If you close the tab, click Open ChatGPT in the extension popup.
Opening another tab does not change the connection. Chrome's internal replacement of the connected tab is followed automatically. If either side of the local connection stops responding, it is retired after 60 seconds without replies so Chrome can reconnect; pending commands are not replayed.
Reasoning level
Both chatgpt_ask and review_with_chatgpt accept reasoning_effort: low, medium, high, or xhigh. Choose by task difficulty: low for bounded simple checks, medium for ordinary analysis, high for substantial reviews or debugging, and xhigh for difficult architecture or ambiguous failures. On the supported ChatGPT UI, low maps to Instant (none) and xhigh to Extra High (max). Omission keeps the current verified non-Pro setting. The selected setting remains in the shared tab after the request.
Every request verifies the selected model and reasoning before sending. Pro is never used: Pro, unknown controls/models, or an unconfirmed requested setting stop the request before Send. Results include requested_reasoning_effort and applied_reasoning (effort, actual UI label and raw value); health exposes the observed reasoning. The current supported model entries are Latest, GPT-5.6 Sol, and GPT-5.5. An unsupported UI needs an adapter update rather than a silent fallback.
Update
From the repository folder:
git pull --ff-only
npm run setupSetup preserves pairing and updates the plugin and connected extension. Open a new Codex task afterward.
Troubleshoot
npm run doctorChecks plugin activation, extension files, MCP startup, and ChatGPT readiness without sending a chat request. Follow the reported next step.
Requests automatically keep reading the existing response after a lost send acknowledgement or a temporary connection failure. New-chat preparation waits for the page to become ready, and response reading tolerates a page reload during generation. If recovery times out, call chatgpt_result with the original request ID. Original inputs are not needed and the prompt is never resent. Completed results are available even when the browser is offline. An active worker returns its current status immediately; recovery after interruption is limited to 30 seconds per call. To stop recovery, use chatgpt_cancel with that ID. Do not create replacement requests while one remains unresolved.
Requests are recorded before context collection or any browser preparation. A preparation failure remains available through chatgpt_result; replay returns that failure without sending. Start a new ID after fixing a confirmed pre-send failure. A crashed worker's lock is reclaimed automatically on the next operation; live workers retain exclusive ownership of the shared tab. Retrying an original tool call with the same ID and identical inputs is also supported. NOT_FOUND means no local record exists; it is not proof that an older server never sent the message.
Results distinguish send_state (not_sent, unknown, confirmed), worker_active, answer_state, and answer_complete. last_observation reports the browser's last observed generation/completion state and timestamp; it is not a live guarantee. Only answer_complete: true is a completed answer. Cancelled fragments remain partial even when they look like a review. cancel_requested: true alone does not confirm cancellation: if its acknowledgement is lost, retrieve or cancel the same ID again. Follow next_action and conversation_reusable; a cancelled or failed latest request cannot reuse an older completed handle.
If an owned response stalls without a Stop control, the extension reloads that conversation once to recover its state. Cancellation also recovers a missing Stop control by reloading and checking ownership again. Existing drafts are preserved. These repairs apply to the shared extension, including requests from other Codex tasks.
Large inputs are inserted as one escaped text fragment with hard line breaks, then read back exactly before sending. This avoids creating an editor paragraph for every source line. Collected context is limited to 47,900 UTF-8 bytes and 1,200 lines, including file/diff wrappers. Larger context is rejected before touching the browser with CONTEXT_TOO_LARGE; narrow the selected paths. The line limit prevents thousands of short lines from freezing the editor despite fitting the byte limit. Inputs get up to 60 seconds for the editor to accept and send them. Preparation, sending, and response reading still share the request's five-minute limit.
Use chatgpt_preview before large requests to inspect bytes, lines, limits, files, and omitted without accessing Chrome or sending content. For an ask, pass mode: ask, prompt, and optional repo_path/context_paths; for a review, pass mode: review, repo_path, scope, and optional paths/base_ref/question. Preview recollects files and does not reserve them. Explicitly selected ignored paths in working-tree reviews are reported as omitted; supply individual files through chatgpt_ask only when intentionally needed. Credential and generated-file exclusions still apply.
COMMAND_EXPIRED means a delayed command was stopped before execution. Cancel its request to clear any unchanged draft it owns, then start a new request with a new ID. Keep Chrome visible and the PC awake. If a response remains stuck, preserve any draft, reload the connected tab, and retrieve the request with its original ID and inputs.
For development, npm test runs the full suite. npm run test:stability repeats the request recovery, worker crash, bridge restart, extension lifecycle, and offline Chrome DOM tests ten times. These simulated fault checks do not replace testing with a signed-in ChatGPT tab. After building, node scripts/live-smoke.mjs --live --interrupt --bridge-restart sends actual test prompts through the connected account and checks sequential requests, follow-ups, replay, large input, and recovery after terminating an active MCP worker and the local bridge. It saves a content-free result summary under artifacts/. Run it only when the shared ChatGPT tab is free. Set CHAT_MCP_TEST_SERVER to test an installed server bundle instead of dist/main.js.
Issue | Fix |
| Install Node.js or Git, then reopen your terminal. |
PowerShell blocks | Use |
Extension fails to load | Rerun setup and select its printed folder, not the repository's |
Waiting for a connection | Click Open ChatGPT in the extension, sign in, then run doctor. |
Tools missing in Codex | Enable Chat MCP and open a new task. |
Already added the GitHub marketplace? Run setup to prepare the runtime and extension. If you also installed the GitHub plugin, enable only the local plugin printed by setup (normally chat-mcp@personal).
For marketplace registration, use https://github.com/phykn/chat-mcp, ref main, and leave Sparse paths empty. Registration alone does not prepare a working installation.
For unattended setup, use npm run setup -- --non-interactive. It skips windows and prompts. Run npm run doctor to check readiness: exit code 0 means ready; 1 means action is needed.
Source layout
src/core owns request state, persistence, hashing, and shared reasoning definitions. src/browser contains page DOM and reasoning controls; src/adapters/extension.ts connects the request runner to the local bridge. extension owns Chrome tab and connection lifecycle, while src/context collects review inputs. The MCP entrypoint and bridge remain src/main.ts and src/bridge.ts.
Usage notice
For individual development workflows, not bulk collection or multi-user hosting. This project is not affiliated with OpenAI.
Browser automation may conflict with service terms and lead to account restrictions. Only send code and prompts you are authorized to share; exclude credentials and confidential information.
License
MIT. Provided without warranty.
Available Tools
4 toolschatgpt_askA
Delegate design comparisons, difficult debugging, or lengthy analysis to ChatGPT. Pass file paths for local context collection. Call once and wait; retry uncertain outcomes with the same request_id and identical inputs.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| repo_path | No | ||
| request_id | Yes | ||
| context_paths | No | ||
| conversation_handle | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral transparency burden. It meaningfully discloses retry/idempotency semantics ('same request_id and identical inputs') and that file paths are used for local context collection. It stops short of explaining output shape, failure modes, or cancellation, but the provided behavior is genuinely useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler; each sentence contributes a distinct piece of information: purpose, context collection, and retry behavior. The most important decision-relevant content is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema and no annotations, the description covers high-level use cases, local context collection, and retry behavior. However, it omits what the response looks like, how repo_path and conversation_handle affect execution, and how this relates to chatgpt_cancel. It is workable but has notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the five parameters. It partially explains file paths/context collection and request_id's role in retries, but it leaves prompt, repo_path, and conversation_handle semantically unexplained. The compensation is insufficient for the parameter count.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Delegate') and names the resource (ChatGPT), followed by concrete task categories: design comparisons, difficult debugging, and lengthy analysis. It is clear about the tool's purpose but does not explicitly contrast it with sibling tools like review_with_chatgpt or chatgpt_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete scenarios where the tool is appropriate and gives actionable guidance: 'Call once and wait; retry uncertain outcomes with the same request_id and identical inputs.' It does not state when not to use it or point to alternatives, so it lacks explicit exclusions but still gives clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatgpt_cancelA
Cancel this request’s owned generation or clear its unchanged unsent draft. Returns cancel_requested while its worker handles cancellation.
| Name | Required | Description | Default |
|---|---|---|---|
| request_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the return value ('cancel_requested'), the asynchronous nature ('while its worker handles cancellation'), and the precise targets of the operation. It doesn't state irreversibility or failure modes, but for a cancellation tool this is solid coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action verb 'Cancel', and the second sentence adds the return value and async status without fluff. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description covers the action, the scope conditions, the return value, and the async behavior. It omits error cases or what happens when there is no owned generation, but these are minor for this level of complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a bare request_id with length constraints and no description. The description anchors that parameter by saying 'this request's owned generation' and 'unchanged unsent draft', giving the agent meaningful semantic context about what request_id identifies and compensating for the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Cancel' and specifies two concrete actions: canceling an owned generation or clearing an unchanged unsent draft. This is a specific verb+resource pairing that clearly sets it apart from siblings like chatgpt_ask, chatgpt_health, and review_with_chatgpt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly defines the scope of use: only an owned generation or an unchanged unsent draft, which gives clear contextual grounding. It does not name alternatives or when-not-to-use cases, but the sibling tool names make the cancellation context self-evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatgpt_healthARead-only
Check once before the first request. Returns ready and a next_action for setup, connection, or recovery; no chat content is exposed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, indicating a safe read operation. The description adds valuable behavioral context: it returns 'ready' and a 'next_action' for setup/connection/recovery, and it explicitly states that no chat content is exposed. This goes beyond the annotation and clarifies what the agent can expect from the call, covering the tool's behavior without contradicting the read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the most important instruction ('Check once before the first request') and then states the return value and a key limitation. There is no waste or redundancy; every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only health check, the description is largely complete. It tells the agent when to call, what it returns (ready and next_action), and what it does not return (chat content). The only minor gap is the lack of detail on the possible values of 'next_action' or error conditions, but given the tool's simplicity and the presence of readOnlyHint, this is acceptable. The description is sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (trivially). The description does not need to explain any parameters. Baseline for no parameters is 4, and there is nothing to add. The description's mention of return values is separate from parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (check) on a specific resource (health), and clarifies it returns ready and next_action for setup/connection/recovery. It also explicitly notes no chat content is exposed, which differentiates it from sibling chat tools like chatgpt_ask and review_with_chatgpt. The verb and scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear when-to-use instruction: 'Check once before the first request.' This implies it is a one-time preflight check. It does not explicitly mention alternatives or when-not-to-use, but the context of siblings (ask, review, cancel) makes it obvious this is the health/readiness gate, not a chat or operation tool. The guidance is sufficient for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_with_chatgptB
Use first for substantial code reviews and once after non-trivial implementation. Collect the diff and matching file contents directly from repo_path; no need to paste code. Narrow paths for focused reviews. Return actionable findings.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | No | ||
| scope | Yes | ||
| base_ref | No | ||
| question | No | ||
| repo_path | Yes | ||
| request_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations or output schema are present, so the description carries the full transparency burden. It adds useful operational behavior: the tool 'collect[s] the diff and matching file contents directly from repo_path' and the agent should 'not paste code.' It also says the tool will 'return actionable findings.' Still, it does not disclose side effects, external call behavior, latency, failure modes, or output structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four short sentences with no filler. It front-loads the usage rule, then gives an operational note, a scoping tip, and the expected output type. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with zero schema coverage and no output schema, the description is too thin. It covers the overall workflow and timing, but leaves important invocation details undocumented: what each scope value means, what base_ref is for, how question/request_id are used, and what 'actionable findings' look like structurally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the six parameters, but it only maps to two of them: 'repo_path' (collect directly from repo_path) and 'paths' (narrow paths). The key parameters 'scope', 'base_ref', 'question', and 'request_id' are left unexplained, including the meaning of the scope enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as the primary option for 'substantial code reviews' and says it returns 'actionable findings', so an agent can tell it is a review-focused ChatGPT tool. It does not explicitly distinguish it from siblings like chatgpt_ask, but the review-specific language and the name make the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit timing guidance: 'Use first for substantial code reviews and once after non-trivial implementation.' It also advises to 'narrow paths for focused reviews.' However, it never states when not to use it or what alternative tools to prefer, stopping short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
chatgpt_ask - First observed
chatgpt_cancel - First observed
chatgpt_health - First observed
review_with_chatgpt
TDQS
Scored across 4 tools
The tools are mostly distinct: health checks, cancelations, general questions, and code reviews each have clear roles. There is mild overlap between chatgpt_ask and review_with_chatgpt for code-related inquiries, but their descriptions give enough separation.
Naming is readable and consistently lowercase snake_case, but the pattern is mixed: chatgpt_ask, chatgpt_health, and chatgpt_cancel use a chatgpt_ prefix, while review_with_chatgpt places the verb first and the service as a suffix. This is a minor but noticeable inconsistency.
Four tools are well-scoped for a ChatGPT integration: ask, review, health, and cancel cover the core workflow without redundancy or bloat. The count feels appropriate for the server's purpose.
The set covers the essential interaction lifecycle: health check, general ask, code review, and cancellation. Minor gaps exist, such as no explicit session/history inspection tool, but agents can likely accomplish their main goals with these tools.
Maintenance
Related MCP Connectors
Persistent memory for Claude Code and Cursor. Stop re-explaining your project every session.
Code intelligence for LLMs. Analyze, search, and retrieve code from any public git repository.
Connect Google Analytics to ChatGPT. Query GA4 data in plain English and get instant insights.
Shared memory for AI coding agents. Save once, reuse from Cursor, Claude Code, Codex.
Related MCP Servers
- AlicenseCqualityFmaintenanceConnects AI assistants like Claude to the Codex CLI for code analysis, editing, and execution. Supports file references with @ syntax, sandboxed code execution with approval workflows, and structured code changes for automated refactoring and documentation.872 npm178MIT
- AlicenseNot gradedqualityAmaintenanceBridges Cursor IDE Agent with ChatGPT Web for external reasoning, research, and review.8 npmMIT
- FlicenseNot gradedqualityBmaintenanceEnables Codex to consult ChatGPT Web through a dedicated Chrome profile as a read-only reviewer and senior advisor, providing second opinions on code, tests, designs, and proposed fixes.12-
- AlicenseNot gradedqualityCmaintenanceEnables Codex to consult ChatGPT Web through a dedicated Chrome session, acting as a read-only reviewer and senior advisor for code, tests, and designs.MIT