opencode-hermes-mcp
opencode-hermes-mcp
[
(#version-pin-opencode-11821)
Hermes(감독 LLM)와 상주(supervisor LLM)와 OpenCode 서버 사이의 결정적(deterministic) MCP 컨트롤러입니다. 컨트롤러는 LLM이 아닌 상태 머신(state machine)으로, OpenCode의 턴에서 블록하고 질응/권한을 Hermes에게 표시해, 감독 LLM이 판단한 뒤 같은 턴을 재개할 수 있게 합니다.
Architecture
Hermes (LLM) --MCP stdio--> opencode_hermes_mcp.server (FastMCP, 6 tools) --HTTP + SSE--> OpenCode server :4096Layer 1 — Hermes: 상독 LLM입니다.
opencode_run로 코딩 작업을 위임하고, 컨트롤러가needs_agent_input(질은/권한)을 리포트하면 판단합니다.Layer 2 — 이 컨트롤러(
opencode_hermes_mcp/:server.py+controller.h+client.py+models.y): Hermes가 MCP stdio를 통해 생성한 NO-LLM 프로세스입니다. 작업을 제출하고 SSE + REST에서 대기하며, 턴이 완료/오류/입력 요구 상태가 될 때까지 블록합니다. 그 후 슈퍼바이저의 판단을 동일한 OpenCode 턴에 다시 전달합니다(프롬프프트는 다시 제출되지 않습니다).Layer 3 — OpenCode 서버: 상주
opencode serve프로세스(systemd 사용자 서비스opencode-server, 프백 :4096, HTTP 베이스 인증)입니다. LLM은~/.config/opencode/opencode.json에서 구멘하는 지원 프롬바이더(OpenAI 호환 엔드포인트, OpenAI, An Thropic) 중 어떠 것이든 사양할 수 있습니다.
Hermes에 노출된 로그: opencode_run, opencode_answer, opencode_permmision, opencode_abort, opencode_inspect(진단 전용), opencode_sessions.
Related MCP server: cursor-agent-bridge
필요 조건
~/.hermes/config.yaml이 있는 Hermes 설치python3>= 3.11(Hermes 구멘vin용 PyYAML 포함)네트워크 액세스(OpenCode 바이너리 설치,
mcp패키지, LLM 엔드포인트)systemd 사용자 세션(
opencode-server서비스용)
설치(명령 2개)
git clone <repo-url> opencode-hermes-mcp && cd opencode-hermes-mcp
scripts/install.shscripts/install.sh는 설치 위자드(opencode_hermes_mcp/installer.py, Python + rich)의 잇은 퍼입니다. 배너, 번호 매 단계, 스타일된 프롬프트, 진행률, 그리고 프로스를 제버닐니다. 보여줍니다. 위자드 스스로 부트스트래프합니다. repo venv가 없으면(또는 rich·pyyaml·mcp==1.12.4·editable 패째지가 없는 경우) venv를 생성한 뒤 재고동하므로, 기본 python3 >= 3.11만 있으면 됩니다.
설치는 멱등적입니다 — 다시 실행하면 에이미 설처된 것은 건너니다. 고정된 OpenCode 바이너리, venv(mcp==1.12.4가 고정된 opencode_hermes_mcp 패키지), LLM 프로바이더 구멘과 시크릿, 서버 좌격 증명, 지된 보호출 2개, systemd 사용자 서비스, 그리고 ~/.hermes/config.yaml 패치(백업 .bak 유지)를 설치합니다. 마지막으로 헬스 체크(제한 헐는 curl --max-time 3, 마지막 오류 노출)와 python -m opencode_hermes_es.moke_client(tool: surface OK 출력)를 마친습니다.
LLM 프로바이더
설치기는 프로바이더에 구속되지 않습니다. 권장하는 세 가지 프롬바이더를 운지합니다:
Provider | Use | npm 패키지 |
| 모든 OpenA-호환 요청(Unsloth, Ollama, vLLM, Haml-server, ...) — 기본 |
|
| 공식 OpenA API |
|
| 공식 Anthropic API |
|
상호작용 모드: 메뉴에서 프로윕더를 선택하고 프롬프트에 답하세요 — openai-compatible은 base URL + API key + model, openai/anthropic은 API key + model, 그다음 LLM 속도(ロー컬 LLM)라면 slow가 프로바이더 옵션에 timeout:false / headerTimeout:false / chunkTimeout:120000을 추가합니다. fast가 기본값), 그리고 모델 제(컨텍스트/출력, 기본 128000 / 32000)을 지정합니다.
비대화형(--yes): 모든 것을 인자 변수에서 읽습니다. 로컬 OpenAI-호환 로кал(Ollama / vLLM / Unsloth / ...):
OPENCODE_PROVIDER=openai-compatible \
OPENCODE_LLM_BASE_URL=http://127.0.0.1:11434/v1 \
OPENCODE_API_KEY=... \
OPENCODE_LLM_MODEL=qwen3.8-27b \
OPENCODE_LLM_SPEED=slow \
scripts/install.sh --yesOpenAI(클라우드):
OPENCODE_PROVIDER=openai OPENCODE_API_KEY=sk-... OPENCODE_LLM_MODEL=gpt-4o \
scripts/install.sh --yesAnthropic(클라우드):
OPENCODE_PROVIDER=anthropic OPENCODE_API_KEY=sk-ant-... \
OPENCODE_LLM_MODEL=claude-sonnet-4-5 scripts/install.sh --yes플래그: --yes(비대화형, OPENCODE_PROVIDER / OPENCODE_LLM_BASE_URL / OPENCODE_API_KEY / OPENCODE_LLM_MODEL / OPENCODE_LLM_SPED / OPENCODE_CONFEXT_LIMIT / OPENCODE_OUTPUT_LIMIT), --port N(기본값 4096), --skip-binary, --force-config, --dry-run, --skip-verify(최종 헬스 + 스모크 인증 생략 — sandbox/CI에 유용).
UNSLOTH_API_KEY는 OPENCODE_API_KEY의 디프리케이트 폴백으로 계속 지원됩니다(기존 스크립트는 그대로 작동).
설치 후 MCP 서버를 로드하려면 새 Hermes 세션이 필요합니다.
Hermes 통합(수동)
설치기가 ~/.hermes/config.yaml을 대신 패치해 주지만, Hermes의 스킬 레이아웃을 변경할 수 없으므로 의도적으로 Hermes 스킬은 설치하지 않습니다. 패키지에 전체 매뉴얼이 포함되어 있습니다:
docs/hermes-integration.md— MCP가 필요한 이유, 정확한 구성 항목, 수동 통합(직접), 6개 도구, 문지 해결, 실행 제거docs/skill.example.md—~/.hermes/skills/에 넣고 실행할 수 있는 준비된 Hermes 스킬(작업 위임 프로토콜)
사용법
Hermes는 도구를 통해 작업을 위임합니다 — CLI를 수동으로 시작할 필요:
opencode_run(directory, task, agent)— 작업을 제출하고, 허턴이 완료되거나 오류가 나거나 입력을 요구할 때까지 블록합니다. 새 세션을 시작하려면agent가 필요합니다(프로젝트의 기본 에이전트, 예:build,plan또는 프로젝트 특화 agent).도구가
state=needs_agent_input를 반환하면 Hermes가 판단합니다:opencode_answer(정확한 옵션 값 선택) 또는opencode_permission(once/always/reject) — 두 가지 모두 같은 턴을 재개합니다.opencode_abort는 더 진행되지 실행 중지;opencode_sessions는 디렉토리의 허턴을 나열합니다;opencode_inspect는 예외 진단용(작업 실행 중 폴링 금지).
Hermes 측 연결 설정(scripts/install.sh이 ~/.hermes.config.yaml에 작성):
mcp_servers:
opencode:
command: ~/.local/bin/opencode-mcp-launch.sh
enabled: true
timeout: 14400
connect_timeout: 30
supports_parallel_tool_calls: false
timeouts:
tools:
sequential_call: 14400
concurrent_batch: 14400런처는 ~/.config/hermes/opencode-server.json에서 OpenCode 서버 자격 증명을 읽고, repo venv 내에서 python -m opencode_hermes_mcp.server를 실행합니다 — config.yaml에는 시크릿가 없이 유지됩니다.
TUI 부착 헬퍼(OpenCode 실시간 보기)
install.sh는 또한 ~/.local/bin/에 두 개의 헬퍼를 넣습니다(소스: scripts/helpers/):
ocattach <repo-abs> [ses_...] # open the OpenCode TUI on a repo / session
oc-current # attach to the session Hermes is supervising NOWocattach— 상주 서버:4096에 대해 OpenCode TUI(opencode attach)를 엽니다 — 별도 tmux 불필요. 세션 ID를 지정하지 않으면 최신 세션을 열거나 선택 옵션을 제공합니다.oc-current현재 가장 최신의~/.local/state/opencode-hermes-mcp/turn_*.json(컨트롤러의 진행 중 상태 상태)를 읽고 해당 세션에 부착합니다 — Hermes가 OpenCode를 구동 중일 때, 실시간으로 실행중 추론을 보기 위해 사용합니다.
모집 ~/.config/hermes/opencode-server.json에서 서버 자격 증명을 읽습니다(컨트롤러 런처와 동일 소스). 허턴이 활성인 동안 TUI에서 Esc/Ctrl+C를 누르지 마요 — OpenCode 측에서 실행 중인 턴이 제공됩니다.
업그레이드 / 제거
scripts/upgrade.sh # controller only: git pull + venv deps + restart + smoke
scripts/upgrade.sh --binary # install the PINNED OpenCode binary (idempotent) — see "Version pin" below
scripts/uninstall.sh # service, launchers, venv, hermes entry, credentials
scripts/uninstall.sh --purge # + OpenCode provider config + API key secret
scripts/uninstall.sh --purge-binary # + the OpenCode binaryuninstall.sh는 purge flags가 아닐 때 git clone, OpenCode 프로바이더 구멘, API 키 시크릿, 바이너리를 건드지 않습니다.
버전 고정: OpenCode 1.18.21
컨트롤러는 버전 앞에 **1.18.21**이 검증된 OpenCode 1.18.21에 대해서인 (해당 버전의 라이브 /doc(시 위의 endpoint contract에 검증됐)이며, 웹 문서 문서의 아님). 고정 값은 opencode_hermes_mcp/in.txt에 있는 단일 진실 소스입니다(한 줄, v 없이): installer.py와 scripts/upgrade.sh 모두 이 것을 참조하고, 파일이 없거나 비어 있을 때 내장 상수로 폴백합니다(예: pip 설치에서 파일이 코드와 함께 제공되지 않는 경우). install.sh는 바이너리를 그 버전으로 고정하고, upgrade.sh은 기본적으로 바이너리를 업그레이드하지 않습니다.
scripts/upgrade.sh --binary(버전 생략)는 고정된 버전을 설치하며 멱등적입니다(이미 고정 버전이면 no-op). --binary latest는 최신 "bleeding edge"를 명시적으로 선택하는 것입니다. --binary X.Y.Z는 요청한 버전을 합니다. 고정 값이 아닌 다른 버전라면 스크릅트를 경고하며, 반드시 컨트롤러를 재검증한 뒤여 신할 수 있습니다:
.venv/bin/python tests/run_tests.py(모든 검증을 통과해야 합니다; 트는 실서 그러한 in 서버를 상대로 MCP std가 컨트롤러를 구동합니다). 실패하면 다시 고정: scripts/upgrade.sh --binary.
Timeouts
세 개의 독립된 타임아웃이 파이프라인을 제한합니다: 컨트롤러 실행 타임아웃(DEFAULT_RUN_TIMEOUT = 36000 s — 단일 opencode_run/opencode_answer/opencode_permission 시도는 1시간 후 포기), MCP 서버 타임아웃(~/.hermes/config.yaml의 mcp_servers.opencode.timeout= 1440 초,connect_timeout = 30 초), **Hermes** 도구 타임아웃 (timeouts.tools.sequential_call/concurrent_batch` = 1440 초) — 바깖 두 개는 컨트롤러의 4배로 설정되어, 길지만 정상인 턴이 감독 계층에서 종료되지 않습니다.
Development
CONTRIBUTING.md에서 config를 확인하세요. 스모크 테스트* 실행법, 통합 스쥐트 실행법, 기여 규칙을 설명합니다.
파일
파일 | 역할 |
| FastMCP stdio 서버 (6개 툴) |
| 상태 머신: submit / wait / resume / classify |
|
|
| 턴/인터액션의 데이다 헬퍼 |
| no-LLM 스모크 테스트 (도구 소면 + 베이스 호출) |
| 완전한 통합 스트(실제 LLM 턴) |
| setup 위자드(Python + rich; 부트스트래핑 venv) |
| 고정 OpenCode 의 버전(단일 소스, 1줄) |
| 이프사이클 ( |
| TUI 부착 험퍼( |
라이선스
MIT — Copyright (c) 2026 Arthur Hottier.
Available Tools
6 toolsopencode_abortA
Abort the active OpenCode session (or a specific one). Does not require the run lock, so it can stop a stuck run.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses that aborting does not require the run lock and can stop a stuck run, but omits critical side effects: whether the session is permanently terminated, whether in-progress work is lost, or any permission requirements. For a destructive operation like abort, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The primary action is front-loaded in the first sentence, and the lock detail is added as a compact second sentence that explains a key differentiator. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values do not need explanation. The description covers the parameter's semantics and the primary use case. However, it lacks edge-case handling: what happens if there is no active session, if the session is already terminated, or if the abort fails. For a tool with a single optional parameter, this is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, leaving session_id completely undefined in the schema. The description compensates by explaining that a null/omitted session_id targets the active session, while a specific value targets that session. This adds meaningful semantics beyond the bare type information, making the parameter's role clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Abort' with a clear resource ('OpenCode session') and distinguishes between the active session and a specific one via optional session_id. The mention of not requiring the run lock further sets it apart from siblings like opencode_sessions (which lists sessions) and opencode_run (which starts them), making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete use case: stopping a stuck run that cannot be aborted normally because the run lock is held. This gives clear context for when to use the tool, but it does not explicitly name alternative tools for different scenarios (e.g., opencode_sessions for listing or opencode_run for starting). The guidance is useful but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
opencode_answerA
Answer a pending OpenCode question (state=needs_agent_input, kind=question) and keep blocking until the turn completes, errors, or needs input again.
answers: ONE entry per sub-question, in order. Each entry is a string or a list of strings. When a sub-question offers options and does not allow custom answers, each value MUST be an exact option label (the server rejects anything else — this tool validates before posting).The answer is posted to OpenCode and the SAME turn resumes (the prompt is never resubmitted).
If the question is no longer pending (already answered/consumed), an error is returned; the turn may have moved on.
| Name | Required | Description | Default |
|---|---|---|---|
| answers | Yes | ||
| timeout | No | ||
| directory | Yes | ||
| session_id | Yes | ||
| question_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility and does so excellently. It discloses blocking until completion/error/needs-input, that the prompt is never resubmitted, that answers are validated before posting, and error handling for stale questions. This provides substantial behavioral context beyond what any annotation might offer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear lead sentence and bullet points for details. Every sentence provides necessary information without redundancy. The format is easy to parse, front-loading the core purpose and then elaborating on behaviors and constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the moderate complexity and presence of an output schema (which covers return values), the description is quite complete. It explains blocking, resumption, validation, and error conditions. The only minor gap is not explicitly stating what happens on successful completion (e.g., turn continues), but that is implied by 'resumes' and covered by output schema. Adequate for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does a thorough job on the most complex parameter 'answers', explaining sub-question ordering, string/list-of-strings format, and exact label requirement for option-based questions. The other parameters (directory, session_id, question_id, timeout) are self-explanatory from names and standard context, so the description covers the critical semantics sufficiently.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the verb and resource: 'Answer a pending OpenCode question' with specific state and kind qualifications (state=needs_agent_input, kind=question). This clearly distinguishes it from sibling tools like opencode_run or opencode_inspect, which handle different actions. The purpose is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by specifying the condition for calling (pending question) and warns about the error if the question is no longer pending. It also notes the blocking behavior and that the same turn resumes. However, it does not explicitly mention alternatives or when NOT to use this tool, though the context makes it fairly clear it's for answering questions within an OpenCode session. The guidance is adequate but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
opencode_inspectA
DIAGNOSTIC ONLY: one-shot snapshot of a session (status, tree, pending permissions/questions, last assistant text). NEVER use this to poll or monitor a running task — opencode_run / opencode_answer / opencode_permission block until the turn ends; polling wastes tokens and is forbidden. Use only for exceptional diagnostics (after a timeout, or to inspect a session you did not start).
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and largely meets it: it discloses the one-shot, non-polling nature and frames the tool as a safe read-only diagnostic. It doesn't explicitly state the NULL session_id behavior or error case for nonexistent sessions, but the core behavioral profile (non-blocking, diagnostic-only, forbidden for monitoring) is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each functional: the first defines scope, the second forbids polling with rationale, the third specifies allowed use cases. The most important trait (DIAGNOSTIC ONLY) is front-loaded. Slightly long but no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary. The description covers purpose, usage constraints, and behavioral traits thoroughly. The main missing piece is the NULL session_id semantics and what happens for sessions the caller did start or that don't exist — a small but real gap for a diagnostic tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there is only a single optional session_id parameter. The description partially compensates by implying session_id selects which session to inspect ('inspect a session you did not start'), but it never defines the parameter's format or what the default NULL value means (current session vs. most recent). This is a genuine gap given zero schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a clear verb and resource ('one-shot snapshot of a session') and enumerates exactly what the snapshot contains (status, tree, pending permissions/questions, last assistant text). It clearly distinguishes itself from polling/monitoring tools, and the full-caps 'DIAGNOSTIC ONLY' prefix makes its role unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
This is exemplary. It explicitly forbids polling or monitoring, names the sibling tools (opencode_run / opencode_answer / opencode_permission) with the reason those are the correct choice (they block until turn ends), and specifies the only valid use cases: exceptional diagnostics after a timeout or inspecting a session you did not start. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
opencode_permissionA
Decide a pending OpenCode permission (state=needs_agent_input, kind=permission) and keep blocking until the turn completes, errors, or needs input again.
reply: exactly one of 'once' (allow this call), 'always' (allow this pattern for the session), 'reject'. Decide as supervisor: allow normal actions necessary for the delegated task; reject destructive or out-of-scope requests.The decision is posted to OpenCode and the SAME turn resumes (the prompt is never resubmitted).
If the permission is no longer pending, an error is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| reply | Yes | ||
| timeout | No | ||
| directory | Yes | ||
| session_id | Yes | ||
| permission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description thoroughly discloses behavior: it blocks until turn completes/errors/needs input, the same turn resumes (prompt never resubmitted), and an error occurs if the permission is no longer pending. This goes well beyond the schema and gives the agent a clear model of execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with bullet points and front-loads the core purpose. It is concise and avoids fluff, though it could be slightly tightened (e.g., repeating 'keep blocking' in the first sentence and subsequent bullets). Overall, it is efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the tool's blocking and error behavior are explained, the description omits the return format of a successful decision and does not clarify the role of `timeout`. Given the tool's complexity (5 parameters, no annotations, and output schema present but not explained), this leaves important gaps for an agent deciding how long to wait or what to expect in response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the `reply` parameter thoroughly (values and semantics), but gives no meaning for `directory`, `session_id`, `permission_id`, or `timeout`. These are left to inference, which is insufficient for a tool with zero other documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (decide a pending OpenCode permission) with clear context (state=needs_agent_input, kind=permission). It distinguishes itself from siblings by focusing on permission decisions, and includes explicit blocking behavior. The purpose is unambiguous and well-scoped.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear decision policy: allow normal actions, reject destructive/out-of-scope requests, and defines the three reply options. However, it does not explicitly mention when not to use this tool or reference alternatives like opencode_abort or opencode_inspect, leaving some inference to the agent about tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
opencode_runA
Delegate a coding task to OpenCode and block until it completes, errors, or needs input (a question or a permission).
New task: pass
directory+task+agent(a new session is created).agentis REQUIRED for a new session.Continuation / resume: pass
session_id(+taskfor a NEW turn on that session, ortaskis ignored when the turn is still in flight).agentis not needed to resume an in-flight turn (it is taken from the turn's durable state); pass it only when starting a fresh turn on an existing session.agent: the OpenCode agent to run as root. Free string, validated dynamically against the project's live agent list (GET /agent). It MUST be a primary agent of that directory (project-specific primary agents are preferred when they fit the task;buildis the generic implementation agent;planis read-only). Subagents are rejected as root. Do not rely on the server's default_agent: always choose explicitly for a new session.model: optional 'provider/model' override.
RESUME: if session_id is given and that session still has a turn in
flight (busy/retry) — e.g. the controller restarted mid-turn — the prompt
is NOT resubmitted: the wait loop simply resumes on the SAME turn
(task and agent are ignored in that case).
The call blocks until the turn ends. While OpenCode works, NOTHING is polled — the controller watches SSE + REST internally. If OpenCode asks a question or requests a permission, the call returns state='needs_agent_input' (kind='question' or 'permission') with everything needed to decide; answer with opencode_answer / opencode_permission, which resume the SAME turn. Returns the final assistant text + diff on completion.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| agent | No | ||
| model | No | ||
| timeout | No | ||
| directory | Yes | ||
| session_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden, and it excels. It discloses the blocking semantics, that 'NOTHING is polled — the controller watches SSE + REST internally', the needs_agent_input return state with kind='question' or 'permission', the RESUME behavior where 'the prompt is NOT resubmitted' for in-flight turns, and the final output (assistant text + diff). No contradiction with annotations since none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but every sentence earns its place for a 6-parameter stateful tool. It is front-loaded with the core blocking purpose, then uses clear scoping (RESUME: heading in caps, bulleted usage modes, bolded parameter names) that makes dense content scannable. The complexity of the state machine fully justifies the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters, zero schema descriptions, and a complex state machine, the prose is remarkably complete: it covers new-task vs continuation, in-flight resume, root-agent restrictions, the question/permission return path, and the completion output. An output schema exists to carry return-value details, and the description handles everything an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the prose must carry parameter meaning, and it does comprehensively: 'agent' gets a rich treatment (root-only, validated against live list, subagents rejected, build vs plan semantics), 'session_id' gets the full resume/in-flight nuance, 'model' is the 'provider/model' override, and 'directory'+'task' are the new-task pair. The only mild gap is 'timeout', documented only by its schema default of 3600, but this is optional and self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb+resource: 'Delegate a coding task to OpenCode and block until it completes, errors, or needs input.' This clearly distinguishes the tool from its siblings (opencode_answer, opencode_permission, opencode_abort, opencode_inspect, opencode_sessions), which are named as complementary follow-ups rather than alternatives to run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance with exclusions: 'New task: pass directory + task + agent' vs. 'Continuation / resume: pass session_id', including the condition that 'task is ignored when the turn is still in flight' and that 'agent' is not needed to resume an in-flight turn. It also names the answer/permission siblings as the path for resuming a turn in needs_agent_input state. No inference is left to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
opencode_sessionsB
List OpenCode sessions for a directory (to pick a session_id to reuse).
| Name | Required | Description | Default |
|---|---|---|---|
| directory | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only says 'List', which implicitly suggests a read-only operation, but it does not state that explicitly, nor does it mention any side effects, authorization requirements, rate limits, or output format details. For a listing tool this is a minor gap, but the description fails to disclose even basic safety or scope constraints beyond the directory parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise and front-loads the core action ('List OpenCode sessions for a directory'). The purpose hint is added in parentheses without verbosity. There is zero wasted content and the structure makes the tool's intent immediately clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with only one parameter, the description conveys the essential information. However, since there are no annotations and schema descriptions are absent, it would benefit from mentioning prerequisites (e.g., OpenCode must be installed or the directory must exist) or clarifying the output structure. The presence of an output schema mitigates the need to explain return values, but behavioral details remain sparse.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate for the undocumented 'directory' parameter. It does add meaning by explaining that the directory scopes the session listing, which is helpful. However, it does not specify the expected format (path, existence requirements, or any constraints) and offers no example. It partially compensates for the schema gap but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List') and the resource ('OpenCode sessions') with a scoping qualifier ('for a directory'). It also adds a purpose hint ('to pick a session_id to reuse'), which clarifies why an agent would call it. It is distinct from the sibling tool names (run, answer, etc.) without confusion, though it does not name any sibling explicitly, so it loses a point for not explicitly differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: it is meant to help pick a session_id for reuse, which suggests it is called before tools like opencode_run or opencode_answer. However, it does not explicitly state when to use this tool versus alternatives, nor does it mention any conditions or exclusions. The guidance is implied rather than explicit, which is adequate but not strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.4.1- First observed
opencode_abort - First observed
opencode_answer - First observed
opencode_inspect - First observed
opencode_permission - First observed
opencode_run - First observed
opencode_sessions
TDQS
Scored across 6 tools
Each tool targets a distinct phase of the OpenCode lifecycle: running/resuming tasks, answering questions, deciding permissions, aborting, inspecting, and listing sessions. The two input-resolution tools are clearly separated by kind (question vs permission). No meaningful overlap exists.
All tools share the opencode_ prefix, which helps, but the second element mixes verbs (run, answer, abort, inspect) with nouns (permission, sessions). There is no consistent verb_noun pattern, though the names remain readable and predictable enough within the server.
Six tools cover the delegated-agent interaction loop without redundancy. Each tool earns its place, and the set is neither bloated nor too thin for the server's purpose.
The core lifecycle is well covered: start/resume tasks, respond to questions, grant or reject permissions, abort, inspect, and list sessions. The main gap is that agents cannot discover the available agent list through a tool, though generic agents like build/plan provide a usable fallback.
Maintenance
Related MCP Connectors
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Persistent memory and cross-session learning for AI coding assistants (hosted remote MCP).
Adaptive plan/build/review cycles for AI coding assistants, persisted across sessions.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables Claude Code to delegate prompts to an OpenCode agent session for cheaper executor-role work, supporting different providers and session persistence.20,872 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients like Claude Code to delegate coding tasks to the local Cursor Agent CLI, with persistent per-workspace sessions that resume across calls.12 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables inspection and control of a running OpenCode TUI session, including live pane capture, prompt injection, interrupting turns, and blocking waits for session state changes.MIT
- AlicenseNot gradedqualityBmaintenanceEnables external AI supervisors to oversee and steer native Codex through MCP, including thread and turn management, observation, interruption, approval responses, runtime status, and checkpointing.MIT