Skip to main content
Glama

cursor-relay-mcp

A local MCP server built on the official @cursor/sdk. It lets MCP clients delegate a bounded task to an explicitly selected Cursor model while keeping durable, restart-safe run state.

@cursor/sdk is public beta. This package pins exactly 1.0.30 and includes an SDK export contract test. See README.zh-CN.md for the full guide.

For a tested Codex Desktop installation workflow on Windows, including personal marketplace layout, CLI fallback, sandbox compatibility, cache refresh, and a Grok 4.6 smoke test, see CODEX_INSTALL.zh-CN.md.

Updating another computer: Pull, build, reinstall, then verify native tools and a real read/write task in a new Codex task. Preserve active runs and local state. See the upgrade checklist.

Codex built-in MCP contract

The Codex plugin manifest references the packaged MCP declaration with "mcpServers": "./.mcp.json". The packaged .mcp.json must start node ./dist/index.js with cwd: ".". Codex resolves that working directory against the installed plugin version root. Do not hard-code a development checkout or a versioned Codex cache path, duplicate the same server in user config.toml, or put CURSOR_API_KEY in the MCP declaration.

After installing or reinstalling the plugin, create a new Codex task. Codex loads the bundled Skill and starts the built-in MCP automatically. It sends only the workspace path, targetLocations (files, directories, or line locations), a bounded task scope, model, permission, and idempotency data. The Cursor Agent reads the authorized workspace itself. Both read-only and workspace-write reject source text, code fences, file contents, and diffs in MCP arguments. An explicit request for Cursor to review a current or named workspace authorizes in-scope reading, not secrets, unrelated paths, or edits. An explicit request to modify or fix must map to workspace-write, not be silently downgraded to analysis. Outside the static allowlist, authorize_workspace issues a reusable capability bound to the current Codex conversation and exact workspace; it supports both read-only and workspace-write.

The plugin ships its Codex Skill inside the original package at skills/delegate-to-cursor-agent/SKILL.md. The manifest declares "skills": "./skills/", and the npm package includes skills/, so installing or reinstalling the plugin makes Codex load the Skill automatically in newly created tasks; no install-time Skill generation is needed.

Related MCP server: LinkedRun

Quick start

Requirements: Git, Node.js >=22.13, and a Cursor account. The recommended authentication method is the official Cursor.auth.login() stored login.

git clone https://github.com/tonytanglab/cursor-relay-mcp.git
cd cursor-relay-mcp
npm install

# Opens the system default browser and stores an official SDK login for this OS user.
node --input-type=module --eval 'import { Cursor } from "@cursor/sdk"; await Cursor.auth.login({ apiKeyName: "cursor-relay-mcp" })'

npm run build
$env:CURSOR_RELAY_WORKSPACE_ROOTS = "D:\app\git"
node .\dist\index.js

Run the login once on each computer and OS user account that starts the MCP server. The official SDK opens the system default browser, mints a named, expiring and revocable API key, and stores it in its official credential store; the relay never reads or returns the key value. Check the login without exposing credentials:

node --input-type=module --eval 'import { Cursor } from "@cursor/sdk"; console.log((await Cursor.auth.status()).status)'

logged-in means the MCP process can use stored login. CURSOR_API_KEY remains an optional alternative for automation, but do not put it in .mcp.json, shell history, logs, or the repository. When using stored login, omit CURSOR_API_KEY entirely instead of setting it to an empty string.

The normal tool flow is doctor → list_models → start_run → repeated wait_run calls until terminal=true. The SDK stored login is independent of the Cursor desktop login; after explicit user confirmation, reauthenticate_cursor can replace a mismatched SDK login without exposing its API key. Runs are idempotent and persisted; the owning executor enforces its total timeout. New runs use workspace-scoped SQLite, while legacy JSONL remains read-only. Reattaching after a restart observes persisted events; it does not restart execution. Stop unconditional polling when needsAttention=true, and do not automatically cancel or resubmit an unknown run.

For start_run and reply_run, task means review/implementation scope and acceptance requirements, never file contents. Use targetLocations for workspace-relative files, directories, or line locations. If a reply omits them, it inherits the parent locations. The Relay builds the Cursor instruction so the agent reads the authorized workspace directly; authorization does not relax this contract for either read-only or read/write runs.

doctor reports the effective defaultTimeoutMs and maxTimeoutMs. Ordinary repository work should normally omit timeoutMs and use the configured default (24 hours by default and also the hard maximum). Callers may request a shorter explicit budget when appropriate, but cannot raise a run above 24 hours. A wait_run timeout is only a polling slice. Retryable SDK reconnects are returned as connection.state=reconnecting and remain non-terminal. While status and events show healthy progress, callers should keep waiting within the run budget instead of cancelling or creating short continuation runs. If no event is persisted for 10 minutes, wait_run returns needsAttention=true and mustCallAgain=false; the run remains non-terminal. The response includes run.activity with the last event time and silence duration. Check the original run and any external process before taking action. Silence alone does not prove that the SDK executor stopped, and must not trigger automatic cancellation or a replacement paid run. The progress page shows the warning while continuing to refresh snapshots so later activity or a terminal result remains visible.

After start_run or reply_run, call open_run and share its clickable progressUrl. This read-only loopback page bypasses MCP App sandbox failures. Links are bound to one run, expire after 24 hours, and stop working when their MCP process exits; call open_run again with the original run ID, never resubmit the task to fix a display error. Do not share these capability links externally. view_run remains an optional embedded panel. Data tools no longer attach a widget, avoiding redundant sandbox frames. Both viewers use read_run_progress for persisted snapshots and the latest 200 incremental events; viewing never attaches to the SDK, settles a timeout, or mutates a run. A persisted running status is not proof of liveness: callers must still monitor with wait_run. The panel retries read failures, shows recovery errors, and releases timers and old events. No network fonts, scripts, or external services are used.

The pinned public Cursor SDK has no operation for injecting a new instruction into an active local Agent run. The relay therefore reports doctor.capabilities.activeRunSteering=false and never treats an internal event append as successful steering. A caller must not cancel merely to redirect: it should keep observing and use reply_run after terminal state. Cancellation is reserved for an explicit stop request or a concrete safety boundary.

Security defaults are fail-closed: unattended runs require the static workspace allowlist. When a user explicitly authorizes Cursor Relay for a workspace in the current conversation, authorize_workspace can issue a reusable read-only or workspace-write capability. The token is never persisted and is bound to the real path, granted permission ceiling, and MCP task/session when available. A workspace-write capability can also run read-only tasks; a read-only capability cannot be elevated. It expires when that conversation scope or MCP process ends. Permissions otherwise default to read-only, the Cursor sandbox is enabled by default on supported non-Windows hosts, and only project settings are loaded. Windows defaults the SDK sandbox off because the current local runtime reports it as unsupported; the read-only tool allowlist remains enforced. Both normal presets expose Cursor's official webSearch and webFetch tools so Cursor can decide whether network research is relevant. danger-full-access still requires the static allowlist plus both CURSOR_RELAY_ENABLE_DANGER_FULL_ACCESS=true at server startup and confirmedDangerousPermission=true on the request. Leave the server switch off for normal use.

CURSOR_RELAY_READ_ONLY_SANDBOX_ENABLED can explicitly disable the sandbox on supported non-Windows hosts. Windows always clamps this setting off because the current Cursor SDK local runtime does not support that sandbox path. This switch applies only to the read-only preset; its public tool allowlist remains restricted to read, grep, glob, ls, webSearch, and webFetch by default. The workspace-write sandbox follows the same platform compatibility rule: it is enabled by default on supported non-Windows hosts and clamped off on Windows, while the exact workspace authorization and disallowed-tool list remain enforced.

delete, task, mcp, and generateImage are controlled by the Codex main process rather than being permanently unavailable. start_run and reply_run accept codexAllowedTools; omitted reply values inherit the parent policy and an empty array revokes it. Read-only runs may additionally allow only generateImage; the other controlled tools require workspace-write. Ordinary, reversible actions already covered by the user's request need no separate prompt. Human confirmation is reserved for materially destructive, broad, irreversible, out-of-scope, externally consequential, or danger-full-access actions.

CURSOR_RELAY_SETTING_SOURCES is an optional comma-separated list restricted to the public project, team, and mdm setting layers. It defaults to project; user, plugins, and all are deliberately rejected to avoid ambient or recursive MCP behavior.

Run summaries omit streamed events and report eventCount; use read_events for event pages. Event data larger than 8 KiB is replaced with explicit truncation metadata. Redaction is based on sensitive field names and is not a general secret scanner, so use a private state directory and avoid secrets in prompts.

The relay uses only public exports from the pinned SDK: model discovery, Agent.create/resume/listRuns/getRun, Run.stream/wait/cancel, and official authentication status. It does not inspect Cursor IDE state or private endpoints. Restarting the Cursor IDE is not required for local SDK runs.

Verification

npm run format:check
npm run lint
npm run typecheck
npm test
npm run build
npm run test:mcp
npm run check:package
npm run test:sdk-contract

The default tests do not call the real Cursor API or modify a real workspace.

Available Tools

14 tools
authorize_workspaceA

仅当用户在当前对话明确要求 Cursor Relay 读取或修改该工作区时调用。签发绑定当前 MCP 对话与精确工作区的可复用授权;read-only 与 workspace-write 均只传位置和范围,禁止传源码正文,由 Cursor 自行读取;不授予危险权限,进程结束即失效。

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceYes
permissionNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety triad (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds real behavioral context: the authorization is reusable, expires when the process ends, does not grant dangerous permissions, and must not receive source-code content. It stops short of describing failure modes or how to revoke/inspect the granted authorization.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the critical gating condition before the behavioral details. Every clause carries information, though the density makes it slightly harder to scan than a two-sentence split would.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with no output schema and thin annotations, the description adequately covers calling condition, scope, lifetime, and the forbidden payload. It does not state what the issued authorization returns or how the agent subsequently consumes it, which is a modest remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate: it explains that both read-only and workspace-write pass only location and scope and never source text, which clarifies the purpose of the `permission` enum values and constrains `workspace`. It still does not define the expected format of `workspace` (path vs. identifier) or the default when `permission` is omitted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (issue/bind authorization) and resource (a precise workspace) and clarifies it binds the current MCP conversation to that workspace. An agent can distinguish it from siblings like reauthenticate_cursor or doctor, though the description never names those siblings explicitly to sharpen the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"仅当用户在当前对话明确要求...时调用" gives a clear, explicit gating condition for when to invoke the tool. It does not name alternative tools or state when *not* to use it beyond the user-request precondition, but the precondition is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_runC
DestructiveIdempotent

取消仍在执行的 Cursor SDK 运行;重复取消具有稳定结果。

ParametersJSON Schema
NameRequiredDescriptionDefault
relayRunIdYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true. The description's note about 'stable results on repeated cancellation' merely restates the idempotency hint without adding new behavioral context, such as what happens to the run, side effects, or permissions required. It adds no value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no fluff. It states the core action and the idempotency property efficiently. The structure is front-loaded with the primary purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (one parameter, no output schema), the description is minimal. It omits key contextual details like the effect on the run lifecycle (e.g., will it stop emitting events?), whether cancellation is reversible, or any return value. An agent would have to infer behavior from the annotations alone, making the description insufficient for full understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the parameter 'relayRunId' is not documented in the schema, and the description does not explain it. Although the name is somewhat self-explanatory, the description fails to compensate for the lack of schema documentation, providing no clarification on its format, purpose, or expected values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'cancel a still-running Cursor SDK run', specifying the verb (cancel) and resource (run). It is distinct from sibling tools like start_run, get_run, etc., though it does not explicitly name alternatives. The purpose is unambiguous and sufficiently specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for runs that are still executing ('仍在执行'), but it provides no explicit guidance on when to use it versus alternatives like wait_run or how it relates to other lifecycle operations. No exclusions or conditions are given beyond the implicit 'still running' state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

doctorA
Read-onlyIdempotent

检查 Cursor Relay 配置、认证与持久目录;不调用 Cursor 模型。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds valuable behavioral context by specifying the exact items checked (config, auth, persistence) and explicitly stating it does not invoke the Cursor model, which goes beyond the annotation hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence with no extraneous words. It front-loads the core purpose and includes an explicit exclusion, making it concise and highly scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only diagnostic tool, the description covers purpose, scope, and the explicit non-action. It lacks detail on the output format, but given the simple nature and no output schema, this is a minor gap. The description is sufficiently complete for an agent to decide when to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description provides context on what the tool inspects without needing to explain parameters, which is appropriate for this parameterless diagnostic tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks Cursor Relay configuration, authentication, and persistence directory, and explicitly notes it does not call the Cursor model. This distinguishes it from sibling run-management tools that interact with models, providing a specific verb and resource scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a diagnostic role by listing what it checks and what it avoids, but it does not explicitly state when to use this tool versus alternatives. It lacks guidance on pre-flight checks or when to prefer it over other operations, leaving usage context to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_runB
Read-onlyIdempotent

读取持久运行并重新附加事件观察;本地执行器退出后不会自动重启模型,execution.state=unknown 时需诊断。

ParametersJSON Schema
NameRequiredDescriptionDefault
relayRunIdYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds genuine behavioral context beyond that: it discloses that the local executor exiting does not auto-restart the model and that an unknown state signals a diagnostic situation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the core action front-loaded and the caveat trailing. No wasted words, though cramming three distinct ideas (read, re-attach, diagnose) into one clause slightly muddies the structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should convey more about the run state it returns, but it only gestures at execution.state=unknown. Annotations cover the safety profile, but the missing parameter documentation and thin return guidance leave gaps for a diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter relayRunId is undocumented in both schema and description. The description never explains what the run identifier is, where to obtain it, or its format, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (读取) and resource (持久运行), plus the secondary effect of re-attaching event observation. This distinguishes it from list_runs and read_events, though it does not clearly separate it from close siblings like view_run and open_run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies the read-and-observe context and notes that execution.state=unknown warrants diagnosis, which is weak guidance on when to use it. It never names an alternative (view_run, open_run, read_run_progress) or states a when-not condition, so routing among the many run-observation siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA
Read-onlyIdempotent

从当前 Cursor 账户发现可用模型、别名与参数;start_run 前应调用。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly, idempotent, and non-destructive hints. The description adds context about using the current account and the ordering requirement relative to start_run, which is useful behavioral information beyond the annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that front-loads the purpose and usage. Every word earns its place; there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, zero-parameter, read-only listing tool, the description adequately specifies what it returns (models, aliases, parameters) and when to call it. With no output schema, it does not explain return structure in detail, but the listed items suffice for an agent to understand the outcome.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description correctly implies no inputs are needed; it adds no parameter-specific information because there are none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('discover') on a specific resource ('available models, aliases, and parameters') from the current Cursor account. It also distinguishes itself from siblings by explicitly referencing start_run, making the tool's role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit timing guidance ('should be called before start_run'), establishing when to use it. It does not mention when not to use it or alternatives, but the context is clear and the directive is direct.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runsB
Read-onlyIdempotent

按创建时间倒序列出 Relay 持久运行。

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

注释已声明readOnlyHint=true和destructiveHint=false,覆盖了安全性。描述额外提到排序方式(创建时间倒序),这是注释未提供的。但未提及返回格式、分页或任何限制,这些对于此类工具是常见的。由于注释已提供基本安全信息,描述本身无需重复,但补充的价值有限。

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

描述仅有一句话,简洁明了,没有任何冗余内容。它直接陈述了核心功能和排序方式,结构清晰。虽然省略了部分细节,但就所包含的信息而言,做到了精炼。

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

该工具参数简单(单可选limit),注释齐全,且无输出schema要求。描述提供了排序行为,但未提及如何设置limit或默认行为。对于这样一个基本工具,描述基本足够,但缺少参数使用说明使上下文略有缺口。

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

输入schema中只有一个可选参数limit,带有最小/最大约束,但schema描述覆盖率为0%。描述完全未提及limit参数的含义或用法,因此代理必须自行推断。虽然'limit'直观代表数量限制,但描述未明确说明,使得参数语义不清晰。

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

描述明确说明了工具的功能:按创建时间倒序列出运行记录。动词'列出'和资源'运行'清晰,且排序方式提供了具体行为。虽然未与其他兄弟工具明确区分,但该描述足以让代理理解基本用途。

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

描述未提供任何关于何时使用该工具与替代工具(如get_run、cancel_run)的明确指导。代理无法从描述中得知适用的场景或与其他列表类工具的区别,仅能通过名称推断。

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_runA
Read-onlyIdempotent

获取现有任务的本机只读进度链接,向用户展示可点击链接;不依赖 Codex MCP 沙箱。start_run/reply_run 成功后调用;沙箱报错时也用它查看原任务,禁止因此重复提交。

ParametersJSON Schema
NameRequiredDescriptionDefault
relayRunIdYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: the link is a local read-only URL, it does not depend on the Codex MCP sandbox, and calling it should never trigger a resubmission. It stops short of describing link lifetime or any rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose (obtain the progress link) before the usage conditions, and every clause carries information. It is a bit dense with three semicolon-joined ideas in one sentence, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool whose annotations already cover safety and whose output (a link to hand to the user) is stated in prose, the definition is largely self-sufficient. The only meaningful gap is the undocumented relayRunId, which is the tool's sole input.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter relayRunId is never explained in the description, including where the agent obtains the id (presumably from start_run/reply_run responses). With one undocumented parameter and no compensation in prose, the agent is left guessing about the identifier's origin and format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: obtain the local read-only progress link for an existing task and surface it to the user as a clickable URL. It is clearly not a mutation or event-reading tool. However, it does not differentiate itself from lookalike siblings such as get_run, view_run, or read_run_progress, which an agent must disambiguate on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit triggers and a prohibition: call it after start_run/reply_run succeed, and also call it when the sandbox errors to inspect the original task. It further states the agent must NOT resubmit as a result. This is textbook when-to-use plus when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_eventsC
Read-onlyIdempotent

按递增序号读取已持久化的 Cursor SDK 流事件。事件数量有上限。

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
relayRunIdYes
afterSequenceNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the increasing-order behavior and the fact that the number of events is capped, which goes beyond annotations. However, it does not disclose whether reading consumes events, how pagination works, or the exact nature of the persisted stream.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler. It front-loads the action and the ordering, which is efficient. However, it is under-specified, so conciseness comes at the cost of completeness, but for what it says, there is zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With three parameters, no parameter descriptions in the schema, and no output schema, this tool requires more context. The description omits critical usage details: how to obtain relayRunId, the meaning and effect of afterSequence, how the limit is applied, and what the returned events look like. It is far from complete for an agent to call correctly without additional context or trial-and-error.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the three parameters (limit, relayRunId, afterSequence). It fails to explain any of them. The only hint is 'events have a limit,' which loosely relates to the limit parameter but is not explicit. No mention of what relayRunId identifies or how afterSequence controls incremental reading.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (read), the resource (persisted Cursor SDK stream events), and the ordering (in increasing sequence number). This is specific and distinct from the sibling tools like get_run (run state) or list_runs (list of runs), so an agent can tell it apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention that it should be used after starting a run or polling for streamed events, nor does it exclude cases where get_run or wait_run would be more appropriate. No alternatives are referenced.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_run_progressA
Read-onlyIdempotent

只读指定任务的持久状态与最近最多 200 条增量事件;不调用 SDK、不重连、不收敛超时。快照不是任务存活证明,实际监控仍用 wait_run。

ParametersJSON Schema
NameRequiredDescriptionDefault
relayRunIdYes
afterSequenceNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/non-open-world, but the description adds real behavior beyond them: it does not invoke the SDK, does not reconnect, does not block on timeout convergence, caps events at 200, and warns that the snapshot is not liveness evidence. This is exactly the kind of consequence-level context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense clause front-loads the purpose, then the operational exclusions, then the routing caveat — no filler and nothing repeated from the name or schema. Sized appropriately for a two-parameter read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read tool with no output schema and strong annotations, the description covers behavior and the liveness caveat well, and loosely characterizes what comes back (state plus bounded events). The remaining gap is the absence of any parameter guidance, which is the one thing an agent still needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the two parameters, and it does not: neither relayRunId's meaning/format nor afterSequence's role as an incremental-event cursor is described. The phrase about 'up to 200 incremental events' hints at incremental semantics but never connects it to afterSequence, leaving the agent to guess how to page.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-resource pair (read-only retrieval of a run's persisted state) and bounds the payload explicitly ('up to the most recent 200 incremental events'). It also differentiates itself from the sibling wait_run by stating that live monitoring belongs to that tool rather than this one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-not rule ('the snapshot is not proof the task is alive; for actual monitoring use wait_run') and names the alternative tool, so an agent can route between a point-in-time snapshot and a blocking wait without inference. It additionally rules out side effects (no SDK call, no reconnect, no timeout convergence).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reauthenticate_cursorA
DestructiveIdempotent

当 Cursor SDK stored login 与用户确认的 Cursor 账户或套餐不一致时,打开系统浏览器重新登录并替换 SDK 本地凭据。不会读取或返回 API key;完成后必须重新调用 list_models 验证权限。

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmedYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly=false, destructive=true, openWorld=true, and idempotent=true. The description adds useful context beyond that: it opens the system browser, replaces SDK local credentials, does not read or return the API key, and requires a validation call afterward. It still omits auth prerequisites and failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded: it states the trigger before the action, then adds security and follow-up constraints in a compact format. No sentence is wasted, though the parameter omission slightly weakens overall density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the trigger, the browser-based action, credential replacement, API key handling, and the required validation step. It leaves the 'confirmed' parameter unexplained and does not contrast with authorize_workspace, but annotations already carry the safety profile, so it is nearly complete for a reauthentication tool without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'confirmed' has no description in the schema (0% coverage) and is never mentioned in the description. The const:true constraint implies a required confirmation, but the definition does not explain why or when to pass it. Since schema coverage is low, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (open system browser to re-login and replace SDK local credentials) and a clear trigger condition. However, it does not explicitly differentiate itself from the sibling authorize_workspace, which also handles an auth-related task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger (stored login inconsistent with user-confirmed Cursor account or plan) and a required follow-up (re-call list_models). It does not state when not to use this tool or name an alternative sibling such as authorize_workspace.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reply_runB
DestructiveIdempotent

在已结束的 Cursor Agent 会话中续接运行并保留 agentId;只传目标位置与任务范围,禁止源码正文,由 Cursor 自行读取。

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。
modelNo
timeoutMsNoCursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。
permissionNo
parentRunIdYes
idempotencyKeyYes
targetLocationsNo工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。
codexAllowedToolsNo本次续接由 Codex 主进程决定的额外工具放行;省略时继承父运行,传空数组可撤销。
workspaceApprovalTokenNo
confirmedDangerousPermissionNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, idempotent=true, openWorld=true, so the safety profile is partly covered. The description adds two useful facts: the agentId is retained and source code must not be inlined (Cursor reads it). It says nothing about permissions, approval tokens, or what a continuation actually mutates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence, front-loaded with the core action and followed by the key constraint. No filler, though the semicolon-packed style condenses several ideas that could be separated for scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world tool with 10 parameters, nested objects, permission modes, an approval token and a danger-confirmation flag, and no output schema, the description is far too thin. It leaves permission handling, idempotency behavior and danger escalation entirely undescribed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40% across 10 parameters, and the description touches only implicitly on task and targetLocations. Critical parameters (permission, workspaceApprovalToken, confirmedDangerousPermission, codexAllowedTools inheritance, idempotencyKey) get no explanation in either the description or the schema, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: '续接运行' on an '已结束的 Cursor Agent 会话', and names a distinguishing trait (retains agentId). It is clearly separable from start_run/open_run by the 'already-ended session' scope, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a condition of use (only for already-ended sessions) and a content rule (pass only locations/scope, never source code). It does not, however, say when to prefer a sibling like start_run or open_run, nor state any exclusion beyond the source-code prohibition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_runB
DestructiveIdempotent

在允许的本地工作区启动持久 Cursor Agent 运行。read-only 与 workspace-write 均只接受工作区、目标位置和任务范围,禁止嵌入源码正文;Cursor 自行读取所需文件。必须显式选择模型和幂等键。

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。
modelYes
timeoutMsNoCursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。
workspaceYes
permissionNo
idempotencyKeyYes
targetLocationsNo工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。
codexAllowedToolsNo仅由 Codex 主进程按任务范围决定的额外工具放行;高风险且超出用户既有授权时应先人工确认。
workspaceApprovalTokenNo
confirmedDangerousPermissionNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered. The description usefully adds the 'no embedded source code' scoping rule and the mandatory model/idempotency key, but says nothing about the danger-full-access path, the approval token, or what destruction actually entails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, the core action front-loaded and constraints following. No filler, though it could be organized slightly more clearly around the parameter groups.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world tool with 10 parameters, nested objects and no output schema, the description covers the required fields and a key constraint but omits the escalation/permission machinery (danger-full-access, approval token, dangerous-permission confirmation) an agent needs to call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 40% schema description coverage, the description must compensate and does partially: it explains that workspace/targetLocations/task carry scope and that model and idempotencyKey are mandatory. But permission, timeoutMs, codexAllowedTools, workspaceApprovalToken and confirmedDangerousPermission get no coverage in either place.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource – '启动持久 Cursor Agent 运行' (start a persistent Cursor Agent run) – so the agent knows exactly what the tool does. However, it does not explicitly distinguish this from siblings such as open_run or wait_run, leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives prerequisites (must explicitly select model and idempotency key) and a constraint (read-only/workspace-write only accept workspace, target locations, task scope), which implies usage context. But there is no explicit when-to-use / when-not-to-use guidance versus the many sibling run tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

view_runA
Read-onlyIdempotent

打开一个只读实时面板,展示指定 Cursor Relay 运行的状态、事件时间线与最终输出;不启动、续接、取消或修改运行。

ParametersJSON Schema
NameRequiredDescriptionDefault
relayRunIdYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine behavior beyond that: this opens an interactive real-time panel and reveals a live event timeline rather than returning a static record, which materially changes how an agent should treat the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence states the purpose, the payload it shows, then the exclusions via a semicolon clause. No filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter view tool with rich annotations and no output schema, the description covers what the agent needs to select and expect. The only real gap is the undocumented parameter, which is minor given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required relayRunId, so the description must compensate and largely does not. 'The specified Cursor Relay run' only restates that some run is targeted — it gives no hint about ID format, where to obtain it, or behavior for an invalid ID.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and deliverable (opens a read-only real-time panel for a run) and enumerates exactly what it surfaces: status, event timeline, final output. It is clearly distinct from mutating siblings, though it doesn't explicitly distinguish itself from read-oriented siblings like get_run, read_events, or read_run_progress.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit negative routing — it does not start, resume, cancel, or modify runs — which tells the agent when NOT to pick it. However, it names no positive alternative (e.g., use get_run for a static snapshot), so the when-to-use side is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_runA
Read-onlyIdempotent

最多等待 30 秒。mustCallAgain=true 时继续轮询;needsAttention=true 时停止无条件轮询并诊断,不能把未知执行状态当作失败或完成。

ParametersJSON Schema
NameRequiredDescriptionDefault
waitMsNo
relayRunIdYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, open-world, non-destructive behavior, so the bar is lower. The description adds the 30-second wait bound, polling decision semantics for `mustCallAgain`/`needsAttention`, and a diagnostic warning about unknown states — useful context beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler. The 30-second cap and polling conditions are front-loaded, making the key constraints immediately visible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description should orient the agent to the return values; it names `mustCallAgain` and `needsAttention` but never defines their semantics or origin. Combined with no `relayRunId` explanation, this leaves gaps, though annotations cover the safety profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain both parameters. It effectively restates the waitMs upper bound as '30 seconds' but says nothing about `relayRunId`, the required run identifier, nor any meaning for `waitMs` beyond the schema maximum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the core action — wait up to 30 seconds and continue polling under specific conditions — which clearly identifies a bounded polling tool. However, it never names the run/resource being waited on and does not distinguish itself from siblings like get_run or view_run, so differentiation relies on the tool name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit conditional instructions for acting on `mustCallAgain` and `needsAttention`, and warns against treating unknown states as failure or completion. It does not say when to choose this tool over `get_run`, `read_run_progress`, or other siblings, so initial-call guidance is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.2
    • Changedauthorize_workspace5 fields changed
      • removedInput schema / properties / idempotencyKey
        Removed value: -{
        -  "maxLength": 200,
        -  "minLength": 8,
        -  "type": "string"
        -}
      • removedInput schema / properties / permission / const
        Removed value: -"read-only"
      • addedInput schema / properties / permission / enum
        Added value: +[
        +  "read-only",
        +  "workspace-write"
        +]
      • removedInput schema / properties / task
        Removed value: -{
        -  "minLength": 1,
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "workspace",
        -  "task",
        -  "idempotencyKey"
        -]New value: +[
        +  "workspace"
        +]
    • Addedopen_run
    • Addedread_run_progress
    • Addedreauthenticate_cursor
    • Changedreply_run7 fields changed
      • addedInput schema / properties / codexAllowedTools
        Added value: +{
        +  "description": "本次续接由 Codex 主进程决定的额外工具放行;省略时继承父运行,传空数组可撤销。",
        +  "items": {
        +    "enum": [
        +      "delete",
        +      "task",
        +      "mcp",
        +      "generateImage"
        +    ],
        +    "type": "string"
        +  },
        +  "maxItems": 4,
        +  "type": "array"
        +}
      • addedInput schema / properties / targetLocations
        Added value: +{
        +  "description": "工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。",
        +  "items": {
        +    "maxLength": 500,
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 100,
        +  "type": "array"
        +}
      • addedInput schema / properties / task / description
        Added value: +"任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。"
      • addedInput schema / properties / task / maxLength
        Added value: +4000
      • addedInput schema / properties / timeoutMs / description
        Added value: +"Cursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。"
      • changedInput schema / properties / timeoutMs / maximum
        Previous value: -9007199254740991New value: +86400000
      • changedInput schema / properties / timeoutMs / minimum
        Previous value: --9007199254740991New value: +1000
    • Changedstart_run7 fields changed
      • addedInput schema / properties / codexAllowedTools
        Added value: +{
        +  "description": "仅由 Codex 主进程按任务范围决定的额外工具放行;高风险且超出用户既有授权时应先人工确认。",
        +  "items": {
        +    "enum": [
        +      "delete",
        +      "task",
        +      "mcp",
        +      "generateImage"
        +    ],
        +    "type": "string"
        +  },
        +  "maxItems": 4,
        +  "type": "array"
        +}
      • addedInput schema / properties / targetLocations
        Added value: +{
        +  "description": "工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。",
        +  "items": {
        +    "maxLength": 500,
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 100,
        +  "type": "array"
        +}
      • addedInput schema / properties / task / description
        Added value: +"任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。"
      • addedInput schema / properties / task / maxLength
        Added value: +4000
      • addedInput schema / properties / timeoutMs / description
        Added value: +"Cursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。"
      • changedInput schema / properties / timeoutMs / maximum
        Previous value: -9007199254740991New value: +86400000
      • changedInput schema / properties / timeoutMs / minimum
        Previous value: --9007199254740991New value: +1000
    • Addedview_run
  2. 10 tool updatesv0.1.0
    • First observedauthorize_workspace
    • First observedcancel_run
    • First observeddoctor
    • First observedget_run
    • First observedlist_models
    • First observedlist_runs
    • First observedread_events
    • First observedreply_run
    • First observedstart_run
    • First observedwait_run

TDQS

B3.4/5.0

Scored across 14 tools

Disambiguation3/5

Several tools target run observation and reading state (open_run, get_run, read_events, read_run_progress, view_run, wait_run), and their boundaries are subtle despite descriptions clarifying snapshot vs. streaming vs. live panel. Core run-control and auth tools are distinct, but the monitoring cluster invites misselection.

Naming Consistency4/5

Almost all tools use snake_case verb_noun-style names such as list_models, start_run, and wait_run, with only doctor as a noun exception. The convention is predictable and readable.

Tool Count4/5

14 tools is reasonable for auth, model discovery, workspace authorization, run lifecycle, and monitoring. The monitoring surface is somewhat heavy, but each tool maps to a defined operation.

Completeness4/5

The surface covers configuration/auth, model discovery, run creation, continuation, cancellation, listing, event reading, and progress monitoring. Explicit cleanup/prune/delete for persisted runs is missing, but the core lifecycle is otherwise covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables agents to submit and manage persistent, dependency-aware task graphs with immutable artifacts, resource reservations, durable event streaming, and retryable process execution over MCP.
    12
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to run declarative agents and DAG workflows as plain tools, with parallel nodes, review loops, and per-run least-privilege sandboxing.
    198 npm
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables durable, least-privilege handoff of bounded tasks from a local producer to a configured MCP worker, with atomic claim, complete, and fail operations under expiring leases.
    MIT