codex-chatgpt-web-mcp
This MCP server lets Codex use an authenticated ChatGPT Web account as a reasoning/chat backend over MCP, with tools to check session status, inspect available models/effort options, and send prompts to get assistant responses.
chatgpt_status: Verify whether the persistent ChatGPT Web session is authenticated and UI-ready; returns flags like
authenticated,uiReady, and currentconversationId.chatgpt_capabilities: Inspect live model picker and reasoning/effort choices exposed by the ChatGPT Web UI (no hard-coded list); returns current selection and available options.
chatgpt_chat: Send a prompt (with optional model, effort, timeout, and conversation ID) to ChatGPT Web and receive the structured assistant response, including response text, byte count, truncation flag, and requested model/effort.
Per README, the server also supports staging input files (diffs, logs, PDFs, images) and structured outputs (code blocks, tables, citations), but these are not exposed as separate MCP tools in the schema—only through the chat tool's prompt. All tools are read-only or task-forbidden; no shell, Git, or repository access is granted to ChatGPT.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codex-chatgpt-web-mcpAsk ChatGPT to review the selected code for edge cases."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
codex-chatgpt-web-mcp
한국어 | English
Use your authenticated ChatGPT Web account as an external reasoning, coding, and review backend for Codex.
Codex keeps ownership of the repository, shell, Git, tests, and patch application. ChatGPT sees only the prompts and files Codex explicitly sends.
Install with Codex
Copy this prompt into a local Codex session:
Install codex-chatgpt-web-mcp on this machine.
Repository:
https://github.com/jiho-symply/codex-chatgpt-web-mcp
Use the normal-user install flow, not the development/source-build flow.
Do not modify files in my current project.
1. Check that Node.js >= 20 and a supported Edge/Chrome/Chromium browser are available.
2. Run:
npx -y codex-chatgpt-web-mcp@latest login
If ChatGPT login, CAPTCHA, or 2FA needs human interaction, stop and ask me to complete it in the opened browser.
3. Register the MCP server with:
codex mcp add chatgpt-web -- npx -y codex-chatgpt-web-mcp@latest mcp
4. Verify registration with:
codex mcp list
5. Do not clone/build the repository unless the documented npx path actually fails.
6. If the current Codex session cannot see the newly added MCP server, tell me to restart Codex.
If anything fails, show me the exact failing command and error instead of guessing.Codex can perform the installation itself if it has local shell permission. You only need to handle interactive ChatGPT login/2FA/CAPTCHA, and possibly restart Codex once after registration.
Related MCP server: MCP-LinkGPT
Manual install
Requirements: Node.js 20+ and a local browser.
Windows: Microsoft Edge or Google Chrome
Linux: Google Chrome or Chromium
# One-time ChatGPT login
npx -y codex-chatgpt-web-mcp@latest login
# Register CGW with Codex
codex mcp add chatgpt-web -- npx -y codex-chatgpt-web-mcp@latest mcpVerify with:
codex mcp listCodex CLI, the ChatGPT/Codex desktop app, and Codex IDE integrations on the same host share the same MCP configuration. UI-only setup and platform details are in docs/installation.md.
Use cases
Second-opinion coding/review — send a diff, implementation, or test result to ChatGPT while Codex remains the orchestrator.
Long reasoning — delegate a difficult analysis and recover the same turn without resending the prompt after timeouts.
File/document analysis — explicitly attach source, logs, PDF, Office documents, CSV/JSON/YAML, screenshots, and images.
Structured outputs — receive code blocks, tables, citations, generated files/images, and other response parts as a structured manifest.
Workspace-isolated context — each local workspace can use its own
CGW-...ChatGPT Project created with Project-only memory.
How it works
Codex ── MCP / stdio ──▶ CGW ── browser ──▶ ChatGPT Web
│ │
├─ repo / shell / Git / tests └─ explicit prompt/files only
└─ validates and applies resultsKey behavior:
one persistent authenticated browser profile, stored outside repositories;
Windows and Linux system-browser auto-detection;
CGW-prefix for newly created workspace Projects;explicit input staging — no arbitrary workspace file reader;
async
send → wait → get_replyflow with idempotent request IDs;structured response extraction for text/code/files/images/tables/citations;
generated assets are staged privately before Codex decides what to do with them;
no shell, Git, patch-apply, arbitrary URL navigation, cookie export, CAPTCHA bypass, or stealth capability is exposed to ChatGPT.
Documentation
Detailed documentation is kept out of this README:
Notes
This is ChatGPT Web browser automation, not the official ChatGPT API.
Initial ChatGPT login is intentionally interactive.
ChatGPT Web UI changes can break selectors; ambiguous UI states fail closed.
Content explicitly uploaded through CGW is sent to the user's ChatGPT account and is subject to ChatGPT retention/settings.
ChatGPT output is untrusted; Codex should validate code and files before use.
License
Available Tools
3 toolschatgpt_capabilitiesChatGPT Web capabilitiesARead-only
Inspect live model and reasoning/effort choices visible to the signed-in ChatGPT account. The web UI is the source of truth; no model list is hard-coded.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| modelPicker | Yes | |
| effortPicker | Yes | |
| flattenedPicker | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers safety, and the description adds meaningful behavioral context: the data is live, account-dependent, and sourced from the web UI rather than a hard-coded list. This helps an agent understand that results may change and should not be assumed static.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The main action is front-loaded, and the second sentence adds an important caveat about the source of truth without unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only introspection tool with an output schema and readOnly annotation, the description fully covers what the agent needs: purpose, source of truth, and dynamic nature. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parametershare, so there is no parameter burden on the description. Schema coverage is complete, and the description appropriately focuses on what is being inspected rather than input details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (inspect) and resource (live model and reasoning/effort choices for the signed-in ChatGPT account). This clearly differentiates it from sibling tools like chatgpt_status and chatgpt_chat, which are about status and chat operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when an agent needs the live, account-visible model and reasoning/effort capabilities rather than chat or status. It does not explicitly name alternatives or exclusions, but the context is clear enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatgpt_chatChatGPT Web chatB
Send prompt text to ChatGPT Web and return the assistant response. No repository or execution capability is granted to ChatGPT. The caller is responsible for selecting minimal context and validating returned code before use.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| effort | No | ||
| prompt | Yes | ||
| timeout_ms | No | ||
| conversation_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| response | Yes | |
| truncated | Yes | |
| responseBytes | Yes | |
| conversationId | Yes | |
| requestedModel | Yes | |
| requestedEffort | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the sparse annotations, the description adds useful behavioral caveats: ChatGPT cannot access repositories or execute code, and returned code must be validated. It does not cover authentication, rate limits, or conversation persistence, but it does reveal non-obvious constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action. Each sentence earns its place: the action, the capability boundary, and the caller responsibility.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema presumably covers return values, but the input side is incomplete for a 5-parameter tool: the meaning of effort, model, timeout_ms, and conversation_id is absent. The behavior caveats are helpful, but an agent still lacks enough detail to use optional parameters correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description needed to compensate, but only 'prompt text' clarifies the prompt parameter. 'model', 'effort', 'timeout_ms', and 'conversation_id' are left to be inferred from their names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies the action ('Send prompt text'), the resource ('ChatGPT Web'), and the result ('return the assistant response'), making the tool's function clear. It does not explicitly contrast with chatgpt_status or chatgpt_capabilities, but the distinct action leaves little room for confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used whenever a ChatGPT response is needed and warns that repository/execution capabilities are absent, acting as a when-not-to-use boundary. It never names the sibling tools or states conditions for choosing status/capabilities instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatgpt_statusChatGPT Web statusARead-only
Check whether the persistent ChatGPT Web session is authenticated and usable. Does not expose cookies, tokens, profile files, or workspace data.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| uiReady | Yes | |
| headless | Yes | |
| authenticated | Yes | |
| conversationId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already present, the description adds useful context by guaranteeing that sensitive data is not exposed. It also clarifies that this is a status check rather than a data-access operation, which helps set expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. The primary purpose is front-loaded, and the second sentence adds a valuable privacy clarification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple zero-parameter tool with an output schema and a read-only annotation. The description covers what the tool checks and what it does not expose, which is sufficient for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides full coverage trivially. The description appropriately focuses on the tool's behavior rather than parameter details, which are unnecessary here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: checking whether the persistent ChatGPT Web session is authenticated and usable. It also explicitly distinguishes itself from data-exposing tools by stating it does not expose cookies, tokens, profile files, or workspace data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is reasonably clear: verify session status before relying on the ChatGPT Web session. However, it does not explicitly name alternatives like chatgpt_capabilities or chatgpt_chat, nor does it state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
chatgpt_capabilities - First observed
chatgpt_chat - First observed
chatgpt_status
TDQS
Scored across 3 tools
Each tool serves a distinct purpose: status checks authentication, capabilities inspects model options, and chat sends messages. There is no overlap or ambiguity between them.
All tools follow a consistent 'chatgpt_' prefix with clear noun suffixes (status, capabilities, chat), forming a predictable and uniform naming convention.
With only 3 tools, the server is tightly scoped to its purpose of interacting with ChatGPT Web, and each tool is essential for the core workflow. This is well within the typical range.
The tool surface covers the essential operations: session validation, capability discovery, and message exchange. No critical gaps are apparent for the stated purpose.
Maintenance
Related MCP Connectors
Use your own Mac from ChatGPT, Claude or Codex: files, commands, documents, and a browser.
Undetectable cloud browser sessions for AI agents and scrapers. Navigate, extract, click, captcha.
Persistent memory and cross-session learning for AI coding assistants (hosted remote MCP).
Persistent cross-session memory shared by Codex, Claude Code, ChatGPT, and other AI agents.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables Codex to interact with Gemini Web through browser automation, allowing it to ask scoped questions and analyze public URLs or YouTube videos without a Gemini API key.3MIT
- FlicenseNot gradedqualityBmaintenanceEnables Codex to consult ChatGPT Web through a dedicated Chrome profile as a read-only reviewer and senior advisor, providing second opinions on code, tests, designs, and proposed fixes.12-
- AlicenseNot gradedqualityCmaintenanceEnables Codex to consult ChatGPT Web through a dedicated Chrome session, acting as a read-only reviewer and senior advisor for code, tests, and designs.MIT
- AlicenseBqualityBmaintenanceEnables MCP clients like Codex to operate chatgpt.com through a dedicated persistent browser profile, allowing chat creation, project management, prompt submission, file uploads, and response reading without using the ChatGPT API.33MIT