Skip to main content
Glama
AxelHu
by AxelHu

chatgpt-web-agent

ChatGPT 웹페이지가 OpenClaw Secure MCP Tunnel을 통해 로컬 도구를 사용할 수 있도록 하는 로컬 MCP 글루 레이어.

이것은 원클릭 설치를 목표로 하는 완성품이 아니라 실행 가능한 참조 구현입니다. 주로 실제로 검증된 접근 방식을 공유하기 위한 것이며, 사용자는 Agent를 활용하여 자신의 로컬 환경에 맞게 빠르게 적응할 수 있습니다.

프로젝트 자체의 MCP 인터페이스는 안정적으로 유지되며, 실제 도구는 교체 가능한 LocalToolBackend가 제공합니다. 첫 번째 백엔드는 OpenClaw Plugin SDK를 직접 재사용하며, OpenClaw 소스 코드를 수정하지 않고 파일, Shell, 패치 및 백그라운드 프로세스 도구를 다시 구현하지 않습니다.

현재 상태

P0은 4가지 OpenClaw 도구를 제공합니다:

  • read

  • exec

  • process

  • apply_patch

선택적으로 Google Drive 파일 교환 도구를 활성화할 수 있습니다:

  • drive_list

  • drive_search

  • drive_stat

  • drive_upload

  • drive_download

  • drive_export

  • drive_mkdir

기본적으로 두 개의 읽기 전용 OpenClaw Skills 도구가 활성화됩니다:

  • skills_list(query?, limit?)

  • skill_read(name)

skills_list()는 현재 eligible + model-visible Skill의 간결한 이름 디렉토리를 반환합니다. 자연어 query가 있는 경우 독립적인 QMD collection을 통해 소수의 후보 이름과 설명을 반환할 수 있습니다. skill_read는 Skill 이름/키만 허용하며, 정식 SKILL.md 경로는 항상 실시간 OpenClaw skills.status에서 해석되며 클라이언트가 제공하는 파일 경로를 받지 않습니다. MCP initialize 지침은 클라이언트에게 다음을 추가로 알립니다: 작업이 명백히 로컬 도구, 서비스, 워크플로우 또는 운영 규칙에 의존할 가능성이 있고 현재 컨텍스트가 부족한 경우에만 Skill을 적극적으로 발견하도록 합니다. 일반적인 자체 포함 작업은 Skill을 조회하지 않습니다.

기본적으로 파일 및 패치 도구만 구성된 workspace에 접근할 수 있으며, exec.workdir도 workspace 내에 위치해야 합니다. OpenClaw 내부의 host/security/ask/node/elevated 매개변수는 MCP 클라이언트에 노출되지 않습니다.

신뢰할 수 있는 ChatGPT workspace에만 권한이 부여된 독립 Connector의 경우 CHATGPT_WEB_AGENT_WORKSPACE_ONLY=false를 설정할 수 있으며, 이 경우 workspace는 상대 경로와 기본 cwd의 기준점일 뿐이며 read/apply_patch/exec.workdir은 외부 절대 경로에 접근할 수 있습니다. 이 모드는 보안 샌드박스가 아닙니다.

exec.workdir 경계는 명령 샌드박스가 아닙니다. exec 권한을 얻은 클라이언트는 여전히 명령 텍스트에서 시스템의 다른 위치에 접근할 수 있습니다. Tunnel은 신뢰할 수 있는 ChatGPT workspace에만 권한을 부여해야 하며, 필요에 따라 OpenClaw의 allowlist/approval 정책을 사용해야 합니다.

Related MCP server: agent-mcp-gateway

개발

Node.js 22.22.3 또는 호환되는 OpenClaw Node 버전과 pnpm이 필요합니다.

pnpm install
pnpm check
pnpm smoke

실행

export CHATGPT_WEB_AGENT_WORKSPACE=/path/to/workspace
# 可信独立 Connector 如需把 workspace 仅作为默认工作目录:
# export CHATGPT_WEB_AGENT_WORKSPACE_ONLY=false
pnpm build
node dist/cli.js

서비스는 MCP stdio를 사용하며, 표준 출력은 MCP 프로토콜만 전달합니다.

설정

.env.example을 복사하여 사용 가능한 환경 변수를 확인하세요. 기본 도구 화이트리스트는 다음과 같습니다:

read,exec,process,apply_patch

exec는 기본적으로 allowlist + on-miss를 사용합니다. 명시적으로 덮어쓸 수 있습니다:

export CHATGPT_WEB_AGENT_EXEC_SECURITY=allowlist
export CHATGPT_WEB_AGENT_EXEC_ASK=on-miss

로컬 신뢰할 수 있는 초기 smoke 테스트 시 임시로 다음을 사용할 수 있습니다:

export CHATGPT_WEB_AGENT_EXEC_SECURITY=full
export CHATGPT_WEB_AGENT_EXEC_ASK=off

아키텍처

ChatGPT Web
  → OpenAI Secure MCP Tunnel
  → chatgpt-web-agent MCP Server
  → LocalToolBackend
      → OpenClawBackend
      → SkillsBackend → OpenClaw Gateway (live status)
                      → QMD MCP (optional semantic discovery)
      → NativeBackend / other backend(后续按需)

Skills 백엔드는 capability discovery/read만 수행합니다. QMD는 후보 검색 가속기일 뿐이며, 실시간 OpenClaw 인벤토리는 항상 eligibility, model visibility 및 정식 Skill 경로의 사실 공급원입니다.

Skills semantic discovery

의미론적 카탈로그는 공유 memory index가 아닌 독립적인 QMD named index에 배치하는 것이 좋습니다. QMD 2.5.3의 vector ANN은 먼저 전체 index에서 후보를 가져온 다음 collection filter를 적용합니다. 수십 개의 Skill을 수만 개의 memory 문서에 섞어 넣으면 작은 collection이 전체 라이브러리 후보에 의해 압도될 수 있습니다.

현재 배포에서는 다음을 사용합니다:

local catalog: <workspace>/skills-catalog/
M4 mirror:     ~/qmd-data/skills-chatgpt-web-agent/
QMD index:     skills-chatgpt-web-agent
collection:    skills-chatgpt-web-agent
MCP endpoint:  http://192.168.0.96:8182/mcp

검색은 Qwen3-Embedding-0.6B, vector-only, rerank=false를 사용하며, query expansion / HyDE를 수행하지 않습니다. QMD 히트는 단지 후보일 뿐입니다. 반환 전에 live skills.status와의 교집합을 구합니다. 카탈로그 스키마/인벤토리는 catalogHash를 통해 세대(generation) 무효화를 수행하며, QMD를 사용할 수 없거나 카탈로그가 오래된 경우 자동으로 live names-only catalog로 폴백합니다.

Google Drive

Drive는 선택적 데이터 채널이며, 백그라운드 동기화, 디스크 마운트 또는 전체 디스크 미러링을 수행하지 않습니다. 구현은 Google Drive API v3를 직접 사용하며, MCP는 작고 안정적인 파일 작업 원시 연산만 노출합니다.

기본적으로 Drive의 로컬 업로드/다운로드/내보내기 경로는 다음 위치에만 있을 수 있습니다:

<CHATGPT_WEB_AGENT_WORKSPACE>/exchange

이 제한은 CHATGPT_WEB_AGENT_WORKSPACE_ONLY와 독립적이며, Drive 도구가 임의의 로컬 데이터 유출 채널로 오용될 위험을 줄이기 위한 것입니다. 실제로 필요한 경우 배포자가 CHATGPT_WEB_AGENT_DRIVE_LOCAL_ROOT 및 CHATGPT_WEB_AGENT_DRIVE_LOCAL_ROOT_ONLY를 통해 조정할 수 있습니다.

일회성 OAuth 설정

  1. Google Cloud에서 Drive API를 활성화하고 Desktop OAuth 클라이언트를 생성합니다.

  2. 다운로드한 OAuth JSON을 다음 위치에 저장합니다:

    <workspace>/.credentials/google-drive/credentials.json

    또는 CHATGPT_WEB_AGENT_DRIVE_CREDENTIALS를 설정하여 다른 로컬 개인 경로를 가리킵니다.

  3. 다음을 실행합니다:

    CHATGPT_WEB_AGENT_WORKSPACE=/path/to/workspace pnpm drive:auth

    브라우저 인증이 완료되면 권한이 0600인 authorized-user 토큰이 생성됩니다. OAuth 클라이언트 시크릿과 갱신 토큰은 Git에 커밋해서는 안 되며, MCP를 통해 반환되지 않습니다.

  4. 서비스 시작 시 다음을 설정합니다:

    export CHATGPT_WEB_AGENT_DRIVE_ENABLED=true

Drive 도구의 folderId / fileId는 Drive API ID를 직접 사용합니다. 일반 바이너리 파일은 drive_download를 사용합니다. Google Docs/Sheets/Slides는 drive_export를 사용하여 지정된 MIME 유형으로 내보냅니다.

Available Tools

6 tools
apply_patchapply_patchC

Apply a patch to one or more files using the apply_patch format. The input should include *** Begin Patch and *** End Patch markers.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesPatch content using the *** Begin Patch/End Patch format.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It states the tool applies patches and uses markers, but it does not mention side effects like file modification, potential for partially applied patches, rollback capability, or whether it creates files that don't exist. Key behavioral traits are missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and front-loaded with the core purpose. It uses two sentences effectively, but the repeated mention of 'apply_patch format' is slightly redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool modifies files, it lacks details on safety, rollback, or how partial patches are handled. With no output schema and no annotations, the description should cover failure modes and post-conditions, which it omits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds minor context about the format markers beyond the schema's generic 'Patch content' description. However, it does not explain the patch syntax (e.g., unified diff), allowed operations, or error handling if format is invalid.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool applies a patch to files and specifies the patch format with markers. However, it does not distinguish itself from siblings like 'exec' or 'process', which might also apply changes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like 'read' or 'exec'. The description lacks context on prerequisites, such as whether files must exist or be writable, and does not mention that 'apply_patch' is specifically for applying patch diffs versus directly editing files.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execexecA

Execute shell commands with background continuation for work that starts now. Use yieldMs/background to continue later via process tool. For long-running work started now, rely on automatic completion wake when it is enabled and the command emits output or fails; otherwise use process to confirm completion. Use process whenever you need logs, status, input, or intervention. Use pty=true for TTY-required commands (terminal UIs, coding agents).

ParametersJSON Schema
NameRequiredDescriptionDefault
envNo
ptyNoRun in a pseudo-terminal (PTY) when available (TTY-required CLIs, coding agents)
commandYesShell command to execute
timeoutNoTimeout in seconds (optional, kills process on expiry)
workdirNoWorking directory. Blank/whitespace values are invalid; omit to use the default cwd.
yieldMsNoMilliseconds to wait before backgrounding (default 10000)
backgroundNoRun in background immediately

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explains backgrounding behavior, automatic completion wake, and the necessity of using 'process' for interaction. However, it does not detail security restrictions, output handling, or timeout effects beyond what is in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, front-loaded with the core purpose, followed by concise usage rules. Every sentence contributes unique information with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description should clarify return values. It does not mention what the initial call returns (e.g., immediate output or a handle). However, it covers the backgrounding workflow, including the role of 'process' for logs and status, making the overall completeness high despite the ambiguity about immediate response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high (86%), so baseline is 3. The description adds usage context for 'yieldMs', 'background', and 'pty', but does not significantly elaborate on parameter meaning beyond what the schema already provides. It ties parameters to scenarios but does not introduce new semantic information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Execute shell commands' and elaborates on background continuation, distinguishing itself from sibling tools like 'process' which handles logs/status. It provides a specific verb-resource combination and explicitly differentiates usage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use alternative tool 'process' (for logs, status, input, intervention) and when to enable 'pty=true' (TTY-required commands). It also explains the 'yieldMs/background' mechanism for continuing work later, leaving no ambiguity about context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

processprocessA

Manage running exec sessions for commands already started: list, poll, log, write, send-keys, submit, paste, kill. Use poll/log when you need status, logs, quiet-success confirmation, or completion confirmation when automatic completion wake is unavailable. Use poll/log also for input-wait hints. Use write/send-keys/submit/paste/kill for input or intervention.

ParametersJSON Schema
NameRequiredDescriptionDefault
eofNoClose stdin after write
hexNoHex bytes to send for send-keys
dataNoData to write for write
keysNoKey tokens to send for send-keys
textNoText to paste for paste
limitNoLog length
actionYesProcess action (list|poll|log|write|send-keys|submit|paste|kill|clear|remove)
offsetNoLog offset
literalNoLiteral string for send-keys
timeoutNoFor poll: wait up to this many milliseconds before returning; max 30000 ms, higher values are clamped to 30000
bracketedNoWrap paste in bracketed mode
sessionIdNoSession id for actions other than list

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses that the tool manages running sessions and that poll can wait up to 30000 ms (clamped). However, it doesn't mention that actions modify session state (e.g., writing data, killing), which is reasonably inferred. It also doesn't state that list returns session IDs needed for other actions, though the schema makes that somewhat clear. Very good but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a tight 4-sentence paragraph. The first sentence front-loads all actions and the tool's purpose. The next two sentences give precise when-to-use guidance. The last sentence covers the remaining actions. Every sentence earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 12 parameters, 100% schema coverage, and no output schema, the description fills the behavioral gap well. It explains when to use each action type. The only missing element is a brief note that some actions (e.g., kill) are destructive or irreversible, which would have pushed completeness to 5. Still strong and sufficient for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description does not add meaning beyond schema property descriptions – it simply names the action groups. The schema already describes each property's purpose (e.g., 'Hex bytes to send for send-keys'). The description adds no new parameter details or usage patterns, so it stays at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description lists all eight supported actions (list, poll, log, write, send-keys, submit, paste, kill) and clearly states the tool manages running exec sessions. It even details specific use cases like confirming completion or getting input-wait hints. This fully distinguishes it from siblings like exec (which starts sessions) and read/apple_patch (which operate on files).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent when to use poll/log for status, logs, or input-wait hints, and when to use write/send-keys/submit/paste/kill for input/intervention. It also mentions the fallback when automatic completion wake is unavailable, giving actionable guidance to select among the tool's own actions. No sibling-level exclusions are needed because the tool manages already-started sessions – distinct from siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

readreadA

Read the contents of a file. Supports text files and images (jpg, png, gif, webp, bmp). Images are sent as attachments. For text files, output is truncated to 2000 lines or 50KB (whichever is hit first). Use offset/limit for large files. When you need the full file, continue with offset until complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the file to read (relative or absolute)
limitNoMaximum number of lines to read
offsetNoLine number to start reading from (1-indexed)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully discloses behavioral traits: truncation to 2000 lines or 50KB, offset/limit pagination, and image handling as attachments. This goes well beyond the input schema's parameter descriptions, providing critical operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long, front-loaded with the purpose, and every sentence provides essential information. No unnecessary words or repetition, making it highly efficient for an AI agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (3 parameters, no output schema), the description covers purpose, supported file types, truncation limits, and pagination strategy. It does not explicitly describe the return format for text files, but the information provided is sufficient for most use cases. A minor gap for completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining the intended use of offset/limit for pagination ('Use offset/limit for large files') and clarifying that offset is 1-indexed, which is implicit in the schema but reinforced here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific verb 'read' and resource 'file', and lists supported file types (text and images). It distinguishes from sibling tools like 'exec' (command execution) and 'skill_read' (reading skills) by focusing on general file reading.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use the tool (for reading text and image files) and how to handle large files via offset/limit pagination ('use offset/limit for large files'). It lacks explicit exclusions (e.g., binary files other than images) but the context is sufficient for correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skill_readA

Read one currently eligible/model-visible OpenClaw Skill by its name or skillKey. Use after skills_list identifies a likely match. This tool does not accept filesystem paths.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It mentions 'currently eligible/model-visible' but does not specify behavior on missing skills, read-only guarantee, or potential errors. The description is minimal regarding side effects or failure conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, no redundant wording. Every clause adds value—purpose, usage timing, and a key constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description could clarify what is returned (e.g., skill content, metadata). It does not mention output details, but the tool's purpose is clear enough for selection. Slightly incomplete in describing the full interaction.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'name' is clarified to accept either the skill name or skillKey, and explicitly excludes filesystem paths. This adds significant meaning beyond the bare string type, compensating for the lack of schema-level description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads a skill by name or skillKey, distinguishing it from sibling tools like skills_list (which lists) and read/apply_patch/exec/process (which operate on files or processes). It is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to use after skills_list identifies a likely match, and provides a clear constraint (does not accept filesystem paths). This fully guides when and how to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skills_listA

Discover local OpenClaw Skills available to ChatGPT Web. Pass a natural-language task description in query to get a small semantic top-k with descriptions. Without query, returns the compact names-only live catalog. Only currently eligible and model-visible Skills are surfaced.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum semantic candidates; defaults to 8.
queryNoNatural-language task or capability to discover.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations present, so description carries full burden. It explains the tool surfaces only 'currently eligible and model-visible Skills' and describes two distinct outputs. This is adequate for a read-operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences pack purpose, dual behavior, and constraints with zero waste. Every sentence contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 optional params, no output schema, no annotations), the description covers core functionality, parameter options, and eligibility. It could mention the return format, but is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. Description adds value by explaining that 'query' triggers semantic top-k search with descriptions, while omitting it yields a names-only catalog. This clarifies the parameter's effect beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool discovers local OpenClaw Skills for ChatGPT Web. It distinguishes between query and no-query modes, but doesn't explicitly differentiate from sibling tool 'skill_read'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context on when to pass a query vs. not, but does not mention alternatives or when to avoid using this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedapply_patch
    • First observedexec
    • First observedprocess
    • First observedread
    • First observedskill_read
    • First observedskills_list

TDQS

A4/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a distinct purpose: file reading, patch application, command execution, process management, and skills listing/reading. Descriptions clearly differentiate them.

Naming Consistency4/5

Most tools use verb-like names with a mix of underscore and no-underscore styles (e.g., 'read' vs 'skills_list'). The convention is not fully uniform but remains understandable.

Tool Count5/5

Six tools is appropriate for a coding-agent server, covering core operations without being excessive or sparse.

Completeness5/5

The suite covers file access, modifications, command execution, background process management, and skill discovery, leaving no obvious gaps for common agent workflows.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables ChatGPT Web to securely access local files and run commands via MCP, with optional OpenCode agent mode for autonomous tasks.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Turns ChatGPT web into a local coding agent, enabling file edits, shell commands, Git operations, patches, and process management through 40+ MCP tools.
    MIT