cursor-relay-mcp
Delegate bounded Cursor Agent tasks from MCP clients with persistent, restart-safe run state and explicit workspace permissions.
Check relay configuration, authentication, and persistence directories without calling Cursor (
doctor).Discover available Cursor models, aliases, and parameters before starting a run (
list_models).Issue a short-lived, one-time read-only workspace authorization when the user explicitly authorizes it in the current conversation (
authorize_workspace).Start a durable Cursor Agent run in an allowed local workspace with explicit model, task scope, permission, timeout, and idempotency key (
start_run).Continue an ended Cursor Agent session with a follow-up run while preserving the same
agentId(reply_run).Read current run status and reconnect to the Cursor SDK run after process restarts (
get_run).Wait up to 30 seconds for progress, then keep polling while
terminal=false(wait_run).Cancel an executing Cursor SDK run, with stable idempotent results (
cancel_run).List persisted Relay runs by creation time (
list_runs).Read persisted Cursor SDK stream events by increasing sequence, with limits and truncation (
read_events).Rely on fail-closed defaults: read-only by default, workspace allowlists, sandbox/permission controls, and dangerous-permission opt-ins.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cursor-relay-mcpstart a run to review the recent code changes and suggest improvements"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
cursor-relay-mcp
A local MCP server built on the official @cursor/sdk. It lets MCP clients delegate a bounded task to an explicitly selected Cursor model while keeping durable, restart-safe run state.
@cursor/sdk is public beta. This package pins exactly 1.0.30 and includes an SDK export contract test. See README.zh-CN.md for the full guide.
For a tested Codex Desktop installation workflow on Windows, including personal marketplace layout, CLI fallback, sandbox compatibility, cache refresh, and a Grok 4.6 smoke test, see CODEX_INSTALL.zh-CN.md.
Updating another computer: Pull, build, reinstall, then verify native tools and a real read/write task in a new Codex task. Preserve active runs and local state. See the upgrade checklist.
Codex built-in MCP contract
The Codex plugin manifest references the packaged MCP declaration with
"mcpServers": "./.mcp.json". The packaged .mcp.json must start
node ./dist/index.js with cwd: ".". Codex resolves that working directory
against the installed plugin version root. Do not hard-code a development checkout
or a versioned Codex cache path, duplicate the same server in user config.toml, or
put CURSOR_API_KEY in the MCP declaration.
After installing or reinstalling the plugin, create a new Codex task. Codex loads
the bundled Skill and starts the built-in MCP automatically. It sends only the
workspace path, targetLocations (files, directories, or line locations), a
bounded task scope, model, permission, and idempotency data. The Cursor Agent
reads the authorized workspace itself. Both read-only and workspace-write
reject source text, code fences, file contents, and diffs in MCP arguments. An
explicit request for Cursor to review a current or named
workspace authorizes in-scope reading, not secrets, unrelated paths, or edits. An
explicit request to modify or fix must map to workspace-write, not be silently
downgraded to analysis. Outside the static allowlist, authorize_workspace
issues a reusable capability bound to the current Codex conversation and exact
workspace; it supports both read-only and workspace-write.
The plugin ships its Codex Skill inside the original package at
skills/delegate-to-cursor-agent/SKILL.md. The manifest declares
"skills": "./skills/", and the npm package includes skills/, so installing
or reinstalling the plugin makes Codex load the Skill automatically in newly
created tasks; no install-time Skill generation is needed.
Related MCP server: LinkedRun
Quick start
Requirements: Git, Node.js >=22.13, and a Cursor account. The recommended
authentication method is the official Cursor.auth.login() stored login.
git clone https://github.com/tonytanglab/cursor-relay-mcp.git
cd cursor-relay-mcp
npm install
# Opens the system default browser and stores an official SDK login for this OS user.
node --input-type=module --eval 'import { Cursor } from "@cursor/sdk"; await Cursor.auth.login({ apiKeyName: "cursor-relay-mcp" })'
npm run build
$env:CURSOR_RELAY_WORKSPACE_ROOTS = "D:\app\git"
node .\dist\index.jsRun the login once on each computer and OS user account that starts the MCP server. The official SDK opens the system default browser, mints a named, expiring and revocable API key, and stores it in its official credential store; the relay never reads or returns the key value. Check the login without exposing credentials:
node --input-type=module --eval 'import { Cursor } from "@cursor/sdk"; console.log((await Cursor.auth.status()).status)'logged-in means the MCP process can use stored login. CURSOR_API_KEY remains
an optional alternative for automation, but do not put it in .mcp.json, shell
history, logs, or the repository. When using stored login, omit
CURSOR_API_KEY entirely instead of setting it to an empty string.
The normal tool flow is doctor → list_models → start_run → repeated wait_run calls until terminal=true. The SDK stored login is independent of the Cursor desktop login; after explicit user confirmation, reauthenticate_cursor can replace a mismatched SDK login without exposing its API key. Runs are idempotent and persisted; the owning executor enforces its total timeout. New runs use workspace-scoped SQLite, while legacy JSONL remains read-only. Reattaching after a restart observes persisted events; it does not restart execution. Stop unconditional polling when needsAttention=true, and do not automatically cancel or resubmit an unknown run.
For start_run and reply_run, task means review/implementation scope and
acceptance requirements, never file contents. Use targetLocations for
workspace-relative files, directories, or line locations. If a reply omits them,
it inherits the parent locations. The Relay builds the Cursor instruction so the
agent reads the authorized workspace directly; authorization does not relax this
contract for either read-only or read/write runs.
doctor reports the effective defaultTimeoutMs and maxTimeoutMs. Ordinary
repository work should normally omit timeoutMs and use the configured default
(24 hours by default and also the hard maximum). Callers may request a shorter
explicit budget when appropriate, but cannot raise a run above 24 hours. A
wait_run timeout is only a polling slice. Retryable SDK reconnects are returned as
connection.state=reconnecting and remain non-terminal.
While status and events show healthy progress, callers should keep waiting within
the run budget instead of cancelling or creating short continuation runs.
If no event is persisted for 10 minutes, wait_run returns needsAttention=true
and mustCallAgain=false; the run remains non-terminal. The response includes
run.activity with the last event time and silence duration. Check the original
run and any external process before taking action. Silence alone does not prove
that the SDK executor stopped, and must not trigger automatic cancellation or
a replacement paid run. The progress page shows the warning while continuing
to refresh snapshots so later activity or a terminal result remains visible.
After start_run or reply_run, call open_run and share its clickable
progressUrl. This read-only loopback page bypasses MCP App sandbox failures.
Links are bound to one run, expire after 24 hours, and stop working when their
MCP process exits; call open_run again with the original run ID, never resubmit
the task to fix a display error. Do not share these capability links externally.
view_run remains an optional embedded panel. Data tools no longer attach a
widget, avoiding redundant sandbox frames. Both viewers use read_run_progress
for persisted snapshots and the latest 200 incremental events; viewing never
attaches to the SDK, settles a timeout, or mutates a run. A persisted running
status is not proof of liveness: callers must still monitor with wait_run.
The panel retries read failures, shows recovery errors, and releases timers
and old events. No network fonts, scripts, or external services are used.
The pinned public Cursor SDK has no operation for injecting a new instruction
into an active local Agent run. The relay therefore reports
doctor.capabilities.activeRunSteering=false and never treats an internal event
append as successful steering. A caller must not cancel merely to redirect: it
should keep observing and use reply_run after terminal state. Cancellation is
reserved for an explicit stop request or a concrete safety boundary.
Security defaults are fail-closed: unattended runs require the static workspace
allowlist. When a user explicitly authorizes Cursor Relay for a workspace in the
current conversation, authorize_workspace can issue a reusable read-only or
workspace-write capability. The token is never persisted and is bound to the
real path, granted permission ceiling, and MCP task/session when available. A
workspace-write capability can also run read-only tasks; a read-only capability
cannot be elevated. It expires when that conversation scope or MCP process ends.
Permissions otherwise default to read-only, the Cursor sandbox is enabled by
default on supported non-Windows hosts, and only project settings are loaded.
Windows defaults the SDK sandbox off because the current local runtime reports
it as unsupported; the read-only tool allowlist remains enforced. Both normal
presets expose Cursor's official webSearch and webFetch tools so Cursor can
decide whether network research is relevant.
danger-full-access still requires the static allowlist plus both
CURSOR_RELAY_ENABLE_DANGER_FULL_ACCESS=true at server startup and
confirmedDangerousPermission=true on the request. Leave the server switch off
for normal use.
CURSOR_RELAY_READ_ONLY_SANDBOX_ENABLED can explicitly disable the sandbox on
supported non-Windows hosts. Windows always clamps this setting off because the
current Cursor SDK local runtime does not support that sandbox path. This switch
applies only to the read-only preset; its public tool allowlist remains
restricted to read, grep, glob, ls, webSearch, and webFetch by
default.
The workspace-write sandbox follows the same platform compatibility rule: it is
enabled by default on supported non-Windows hosts and clamped off on Windows,
while the exact workspace authorization and disallowed-tool list remain enforced.
delete, task, mcp, and generateImage are controlled by the Codex main
process rather than being permanently unavailable. start_run and reply_run
accept codexAllowedTools; omitted reply values inherit the parent policy and
an empty array revokes it. Read-only runs may additionally allow only
generateImage; the other controlled tools require workspace-write. Ordinary,
reversible actions already covered by the user's request need no separate prompt.
Human confirmation is reserved for materially destructive, broad, irreversible,
out-of-scope, externally consequential, or danger-full-access actions.
CURSOR_RELAY_SETTING_SOURCES is an optional comma-separated list restricted
to the public project, team, and mdm setting layers. It defaults to
project; user, plugins, and all are deliberately rejected to avoid
ambient or recursive MCP behavior.
Run summaries omit streamed events and report eventCount; use read_events
for event pages. Event data larger than 8 KiB is replaced with explicit
truncation metadata. Redaction is based on sensitive field names and is not a
general secret scanner, so use a private state directory and avoid secrets in
prompts.
The relay uses only public exports from the pinned SDK: model discovery,
Agent.create/resume/listRuns/getRun, Run.stream/wait/cancel, and
official authentication status. It does not inspect Cursor IDE state or private
endpoints. Restarting the Cursor IDE is not required for local SDK runs.
Verification
npm run format:check
npm run lint
npm run typecheck
npm test
npm run build
npm run test:mcp
npm run check:package
npm run test:sdk-contractThe default tests do not call the real Cursor API or modify a real workspace.
Available Tools
14 toolsauthorize_workspaceA
仅当用户在当前对话明确要求 Cursor Relay 读取或修改该工作区时调用。签发绑定当前 MCP 对话与精确工作区的可复用授权;read-only 与 workspace-write 均只传位置和范围,禁止传源码正文,由 Cursor 自行读取;不授予危险权限,进程结束即失效。
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | Yes | ||
| permission | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety triad (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds real behavioral context: the authorization is reusable, expires when the process ends, does not grant dangerous permissions, and must not receive source-code content. It stops short of describing failure modes or how to revoke/inspect the granted authorization.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the critical gating condition before the behavioral details. Every clause carries information, though the density makes it slightly harder to scan than a two-sentence split would.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no output schema and thin annotations, the description adequately covers calling condition, scope, lifetime, and the forbidden payload. It does not state what the issued authorization returns or how the agent subsequently consumes it, which is a modest remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate: it explains that both read-only and workspace-write pass only location and scope and never source text, which clarifies the purpose of the `permission` enum values and constrains `workspace`. It still does not define the expected format of `workspace` (path vs. identifier) or the default when `permission` is omitted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (issue/bind authorization) and resource (a precise workspace) and clarifies it binds the current MCP conversation to that workspace. An agent can distinguish it from siblings like reauthenticate_cursor or doctor, though the description never names those siblings explicitly to sharpen the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"仅当用户在当前对话明确要求...时调用" gives a clear, explicit gating condition for when to invoke the tool. It does not name alternative tools or state when *not* to use it beyond the user-request precondition, but the precondition is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_runCDestructiveIdempotent
取消仍在执行的 Cursor SDK 运行;重复取消具有稳定结果。
| Name | Required | Description | Default |
|---|---|---|---|
| relayRunId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true. The description's note about 'stable results on repeated cancellation' merely restates the idempotency hint without adding new behavioral context, such as what happens to the run, side effects, or permissions required. It adds no value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no fluff. It states the core action and the idempotency property efficiently. The structure is front-loaded with the primary purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (one parameter, no output schema), the description is minimal. It omits key contextual details like the effect on the run lifecycle (e.g., will it stop emitting events?), whether cancellation is reversible, or any return value. An agent would have to infer behavior from the annotations alone, making the description insufficient for full understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the parameter 'relayRunId' is not documented in the schema, and the description does not explain it. Although the name is somewhat self-explanatory, the description fails to compensate for the lack of schema documentation, providing no clarification on its format, purpose, or expected values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'cancel a still-running Cursor SDK run', specifying the verb (cancel) and resource (run). It is distinct from sibling tools like start_run, get_run, etc., though it does not explicitly name alternatives. The purpose is unambiguous and sufficiently specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for runs that are still executing ('仍在执行'), but it provides no explicit guidance on when to use it versus alternatives like wait_run or how it relates to other lifecycle operations. No exclusions or conditions are given beyond the implicit 'still running' state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
doctorARead-onlyIdempotent
检查 Cursor Relay 配置、认证与持久目录;不调用 Cursor 模型。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds valuable behavioral context by specifying the exact items checked (config, auth, persistence) and explicitly stating it does not invoke the Cursor model, which goes beyond the annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with no extraneous words. It front-loads the core purpose and includes an explicit exclusion, making it concise and highly scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only diagnostic tool, the description covers purpose, scope, and the explicit non-action. It lacks detail on the output format, but given the simple nature and no output schema, this is a minor gap. The description is sufficiently complete for an agent to decide when to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description provides context on what the tool inspects without needing to explain parameters, which is appropriate for this parameterless diagnostic tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks Cursor Relay configuration, authentication, and persistence directory, and explicitly notes it does not call the Cursor model. This distinguishes it from sibling run-management tools that interact with models, providing a specific verb and resource scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a diagnostic role by listing what it checks and what it avoids, but it does not explicitly state when to use this tool versus alternatives. It lacks guidance on pre-flight checks or when to prefer it over other operations, leaving usage context to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_runBRead-onlyIdempotent
读取持久运行并重新附加事件观察;本地执行器退出后不会自动重启模型,execution.state=unknown 时需诊断。
| Name | Required | Description | Default |
|---|---|---|---|
| relayRunId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds genuine behavioral context beyond that: it discloses that the local executor exiting does not auto-restart the model and that an unknown state signals a diagnostic situation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core action front-loaded and the caveat trailing. No wasted words, though cramming three distinct ideas (read, re-attach, diagnose) into one clause slightly muddies the structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should convey more about the run state it returns, but it only gestures at execution.state=unknown. Annotations cover the safety profile, but the missing parameter documentation and thin return guidance leave gaps for a diagnostic tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter relayRunId is undocumented in both schema and description. The description never explains what the run identifier is, where to obtain it, or its format, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (读取) and resource (持久运行), plus the secondary effect of re-attaching event observation. This distinguishes it from list_runs and read_events, though it does not clearly separate it from close siblings like view_run and open_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies the read-and-observe context and notes that execution.state=unknown warrants diagnosis, which is weak guidance on when to use it. It never names an alternative (view_run, open_run, read_run_progress) or states a when-not condition, so routing among the many run-observation siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsARead-onlyIdempotent
从当前 Cursor 账户发现可用模型、别名与参数;start_run 前应调用。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, and non-destructive hints. The description adds context about using the current account and the ordering requirement relative to start_run, which is useful behavioral information beyond the annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the purpose and usage. Every word earns its place; there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-parameter, read-only listing tool, the description adequately specifies what it returns (models, aliases, parameters) and when to call it. With no output schema, it does not explain return structure in detail, but the listed items suffice for an agent to understand the outcome.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description correctly implies no inputs are needed; it adds no parameter-specific information because there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('discover') on a specific resource ('available models, aliases, and parameters') from the current Cursor account. It also distinguishes itself from siblings by explicitly referencing start_run, making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit timing guidance ('should be called before start_run'), establishing when to use it. It does not mention when not to use it or alternatives, but the context is clear and the directive is direct.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsBRead-onlyIdempotent
按创建时间倒序列出 Relay 持久运行。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
注释已声明readOnlyHint=true和destructiveHint=false,覆盖了安全性。描述额外提到排序方式(创建时间倒序),这是注释未提供的。但未提及返回格式、分页或任何限制,这些对于此类工具是常见的。由于注释已提供基本安全信息,描述本身无需重复,但补充的价值有限。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
描述仅有一句话,简洁明了,没有任何冗余内容。它直接陈述了核心功能和排序方式,结构清晰。虽然省略了部分细节,但就所包含的信息而言,做到了精炼。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
该工具参数简单(单可选limit),注释齐全,且无输出schema要求。描述提供了排序行为,但未提及如何设置limit或默认行为。对于这样一个基本工具,描述基本足够,但缺少参数使用说明使上下文略有缺口。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
输入schema中只有一个可选参数limit,带有最小/最大约束,但schema描述覆盖率为0%。描述完全未提及limit参数的含义或用法,因此代理必须自行推断。虽然'limit'直观代表数量限制,但描述未明确说明,使得参数语义不清晰。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
描述明确说明了工具的功能:按创建时间倒序列出运行记录。动词'列出'和资源'运行'清晰,且排序方式提供了具体行为。虽然未与其他兄弟工具明确区分,但该描述足以让代理理解基本用途。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
描述未提供任何关于何时使用该工具与替代工具(如get_run、cancel_run)的明确指导。代理无法从描述中得知适用的场景或与其他列表类工具的区别,仅能通过名称推断。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_runARead-onlyIdempotent
获取现有任务的本机只读进度链接,向用户展示可点击链接;不依赖 Codex MCP 沙箱。start_run/reply_run 成功后调用;沙箱报错时也用它查看原任务,禁止因此重复提交。
| Name | Required | Description | Default |
|---|---|---|---|
| relayRunId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: the link is a local read-only URL, it does not depend on the Codex MCP sandbox, and calling it should never trigger a resubmission. It stops short of describing link lifetime or any rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose (obtain the progress link) before the usage conditions, and every clause carries information. It is a bit dense with three semicolon-joined ideas in one sentence, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool whose annotations already cover safety and whose output (a link to hand to the user) is stated in prose, the definition is largely self-sufficient. The only meaningful gap is the undocumented relayRunId, which is the tool's sole input.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter relayRunId is never explained in the description, including where the agent obtains the id (presumably from start_run/reply_run responses). With one undocumented parameter and no compensation in prose, the agent is left guessing about the identifier's origin and format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: obtain the local read-only progress link for an existing task and surface it to the user as a clickable URL. It is clearly not a mutation or event-reading tool. However, it does not differentiate itself from lookalike siblings such as get_run, view_run, or read_run_progress, which an agent must disambiguate on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers and a prohibition: call it after start_run/reply_run succeed, and also call it when the sandbox errors to inspect the original task. It further states the agent must NOT resubmit as a result. This is textbook when-to-use plus when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_eventsCRead-onlyIdempotent
按递增序号读取已持久化的 Cursor SDK 流事件。事件数量有上限。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| relayRunId | Yes | ||
| afterSequence | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the increasing-order behavior and the fact that the number of events is capped, which goes beyond annotations. However, it does not disclose whether reading consumes events, how pagination works, or the exact nature of the persisted stream.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It front-loads the action and the ordering, which is efficient. However, it is under-specified, so conciseness comes at the cost of completeness, but for what it says, there is zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With three parameters, no parameter descriptions in the schema, and no output schema, this tool requires more context. The description omits critical usage details: how to obtain relayRunId, the meaning and effect of afterSequence, how the limit is applied, and what the returned events look like. It is far from complete for an agent to call correctly without additional context or trial-and-error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three parameters (limit, relayRunId, afterSequence). It fails to explain any of them. The only hint is 'events have a limit,' which loosely relates to the limit parameter but is not explicit. No mention of what relayRunId identifies or how afterSequence controls incremental reading.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (read), the resource (persisted Cursor SDK stream events), and the ordering (in increasing sequence number). This is specific and distinct from the sibling tools like get_run (run state) or list_runs (list of runs), so an agent can tell it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention that it should be used after starting a run or polling for streamed events, nor does it exclude cases where get_run or wait_run would be more appropriate. No alternatives are referenced.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_run_progressARead-onlyIdempotent
只读指定任务的持久状态与最近最多 200 条增量事件;不调用 SDK、不重连、不收敛超时。快照不是任务存活证明,实际监控仍用 wait_run。
| Name | Required | Description | Default |
|---|---|---|---|
| relayRunId | Yes | ||
| afterSequence | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/non-open-world, but the description adds real behavior beyond them: it does not invoke the SDK, does not reconnect, does not block on timeout convergence, caps events at 200, and warns that the snapshot is not liveness evidence. This is exactly the kind of consequence-level context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense clause front-loads the purpose, then the operational exclusions, then the routing caveat — no filler and nothing repeated from the name or schema. Sized appropriately for a two-parameter read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read tool with no output schema and strong annotations, the description covers behavior and the liveness caveat well, and loosely characterizes what comes back (state plus bounded events). The remaining gap is the absence of any parameter guidance, which is the one thing an agent still needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the two parameters, and it does not: neither relayRunId's meaning/format nor afterSequence's role as an incremental-event cursor is described. The phrase about 'up to 200 incremental events' hints at incremental semantics but never connects it to afterSequence, leaving the agent to guess how to page.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-resource pair (read-only retrieval of a run's persisted state) and bounds the payload explicitly ('up to the most recent 200 incremental events'). It also differentiates itself from the sibling wait_run by stating that live monitoring belongs to that tool rather than this one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-not rule ('the snapshot is not proof the task is alive; for actual monitoring use wait_run') and names the alternative tool, so an agent can route between a point-in-time snapshot and a blocking wait without inference. It additionally rules out side effects (no SDK call, no reconnect, no timeout convergence).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reauthenticate_cursorADestructiveIdempotent
当 Cursor SDK stored login 与用户确认的 Cursor 账户或套餐不一致时,打开系统浏览器重新登录并替换 SDK 本地凭据。不会读取或返回 API key;完成后必须重新调用 list_models 验证权限。
| Name | Required | Description | Default |
|---|---|---|---|
| confirmed | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, destructive=true, openWorld=true, and idempotent=true. The description adds useful context beyond that: it opens the system browser, replaces SDK local credentials, does not read or return the API key, and requires a validation call afterward. It still omits auth prerequisites and failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: it states the trigger before the action, then adds security and follow-up constraints in a compact format. No sentence is wasted, though the parameter omission slightly weakens overall density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the trigger, the browser-based action, credential replacement, API key handling, and the required validation step. It leaves the 'confirmed' parameter unexplained and does not contrast with authorize_workspace, but annotations already carry the safety profile, so it is nearly complete for a reauthentication tool without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The sole parameter 'confirmed' has no description in the schema (0% coverage) and is never mentioned in the description. The const:true constraint implies a required confirmation, but the definition does not explain why or when to pass it. Since schema coverage is low, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (open system browser to re-login and replace SDK local credentials) and a clear trigger condition. However, it does not explicitly differentiate itself from the sibling authorize_workspace, which also handles an auth-related task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger (stored login inconsistent with user-confirmed Cursor account or plan) and a required follow-up (re-call list_models). It does not state when not to use this tool or name an alternative sibling such as authorize_workspace.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_runBDestructiveIdempotent
在已结束的 Cursor Agent 会话中续接运行并保留 agentId;只传目标位置与任务范围,禁止源码正文,由 Cursor 自行读取。
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | 任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。 | |
| model | No | ||
| timeoutMs | No | Cursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。 | |
| permission | No | ||
| parentRunId | Yes | ||
| idempotencyKey | Yes | ||
| targetLocations | No | 工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。 | |
| codexAllowedTools | No | 本次续接由 Codex 主进程决定的额外工具放行;省略时继承父运行,传空数组可撤销。 | |
| workspaceApprovalToken | No | ||
| confirmedDangerousPermission | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=true, openWorld=true, so the safety profile is partly covered. The description adds two useful facts: the agentId is retained and source code must not be inlined (Cursor reads it). It says nothing about permissions, approval tokens, or what a continuation actually mutates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence, front-loaded with the core action and followed by the key constraint. No filler, though the semicolon-packed style condenses several ideas that could be separated for scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, open-world tool with 10 parameters, nested objects, permission modes, an approval token and a danger-confirmation flag, and no output schema, the description is far too thin. It leaves permission handling, idempotency behavior and danger escalation entirely undescribed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% across 10 parameters, and the description touches only implicitly on task and targetLocations. Critical parameters (permission, workspaceApprovalToken, confirmedDangerousPermission, codexAllowedTools inheritance, idempotencyKey) get no explanation in either the description or the schema, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: '续接运行' on an '已结束的 Cursor Agent 会话', and names a distinguishing trait (retains agentId). It is clearly separable from start_run/open_run by the 'already-ended session' scope, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a condition of use (only for already-ended sessions) and a content rule (pass only locations/scope, never source code). It does not, however, say when to prefer a sibling like start_run or open_run, nor state any exclusion beyond the source-code prohibition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_runBDestructiveIdempotent
在允许的本地工作区启动持久 Cursor Agent 运行。read-only 与 workspace-write 均只接受工作区、目标位置和任务范围,禁止嵌入源码正文;Cursor 自行读取所需文件。必须显式选择模型和幂等键。
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | 任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。 | |
| model | Yes | ||
| timeoutMs | No | Cursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。 | |
| workspace | Yes | ||
| permission | No | ||
| idempotencyKey | Yes | ||
| targetLocations | No | 工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。 | |
| codexAllowedTools | No | 仅由 Codex 主进程按任务范围决定的额外工具放行;高风险且超出用户既有授权时应先人工确认。 | |
| workspaceApprovalToken | No | ||
| confirmedDangerousPermission | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered. The description usefully adds the 'no embedded source code' scoping rule and the mandatory model/idempotency key, but says nothing about the danger-full-access path, the approval token, or what destruction actually entails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, the core action front-loaded and constraints following. No filler, though it could be organized slightly more clearly around the parameter groups.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, open-world tool with 10 parameters, nested objects and no output schema, the description covers the required fields and a key constraint but omits the escalation/permission machinery (danger-full-access, approval token, dangerous-permission confirmation) an agent needs to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 40% schema description coverage, the description must compensate and does partially: it explains that workspace/targetLocations/task carry scope and that model and idempotencyKey are mandatory. But permission, timeoutMs, codexAllowedTools, workspaceApprovalToken and confirmedDangerousPermission get no coverage in either place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource – '启动持久 Cursor Agent 运行' (start a persistent Cursor Agent run) – so the agent knows exactly what the tool does. However, it does not explicitly distinguish this from siblings such as open_run or wait_run, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives prerequisites (must explicitly select model and idempotency key) and a constraint (read-only/workspace-write only accept workspace, target locations, task scope), which implies usage context. But there is no explicit when-to-use / when-not-to-use guidance versus the many sibling run tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
view_runARead-onlyIdempotent
打开一个只读实时面板,展示指定 Cursor Relay 运行的状态、事件时间线与最终输出;不启动、续接、取消或修改运行。
| Name | Required | Description | Default |
|---|---|---|---|
| relayRunId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine behavior beyond that: this opens an interactive real-time panel and reveals a live event timeline rather than returning a static record, which materially changes how an agent should treat the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence states the purpose, the payload it shows, then the exclusions via a semicolon clause. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter view tool with rich annotations and no output schema, the description covers what the agent needs to select and expect. The only real gap is the undocumented parameter, which is minor given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required relayRunId, so the description must compensate and largely does not. 'The specified Cursor Relay run' only restates that some run is targeted — it gives no hint about ID format, where to obtain it, or behavior for an invalid ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and deliverable (opens a read-only real-time panel for a run) and enumerates exactly what it surfaces: status, event timeline, final output. It is clearly distinct from mutating siblings, though it doesn't explicitly distinguish itself from read-oriented siblings like get_run, read_events, or read_run_progress.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit negative routing — it does not start, resume, cancel, or modify runs — which tells the agent when NOT to pick it. However, it names no positive alternative (e.g., use get_run for a static snapshot), so the when-to-use side is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_runARead-onlyIdempotent
最多等待 30 秒。mustCallAgain=true 时继续轮询;needsAttention=true 时停止无条件轮询并诊断,不能把未知执行状态当作失败或完成。
| Name | Required | Description | Default |
|---|---|---|---|
| waitMs | No | ||
| relayRunId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, open-world, non-destructive behavior, so the bar is lower. The description adds the 30-second wait bound, polling decision semantics for `mustCallAgain`/`needsAttention`, and a diagnostic warning about unknown states — useful context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler. The 30-second cap and polling conditions are front-loaded, making the key constraints immediately visible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description should orient the agent to the return values; it names `mustCallAgain` and `needsAttention` but never defines their semantics or origin. Combined with no `relayRunId` explanation, this leaves gaps, though annotations cover the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain both parameters. It effectively restates the waitMs upper bound as '30 seconds' but says nothing about `relayRunId`, the required run identifier, nor any meaning for `waitMs` beyond the schema maximum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the core action — wait up to 30 seconds and continue polling under specific conditions — which clearly identifies a bounded polling tool. However, it never names the run/resource being waited on and does not distinguish itself from siblings like get_run or view_run, so differentiation relies on the tool name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit conditional instructions for acting on `mustCallAgain` and `needsAttention`, and warns against treating unknown states as failure or completion. It does not say when to choose this tool over `get_run`, `read_run_progress`, or other siblings, so initial-call guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.2- Changed
authorize_workspace5 fields changed- removed
Input schema / properties / idempotencyKeyRemoved value: -{ - "maxLength": 200, - "minLength": 8, - "type": "string" -} - removed
Input schema / properties / permission / constRemoved value: -"read-only" - added
Input schema / properties / permission / enumAdded value: +[ + "read-only", + "workspace-write" +] - removed
Input schema / properties / taskRemoved value: -{ - "minLength": 1, - "type": "string" -} - changed
Input schema / requiredPrevious value: -[ - "workspace", - "task", - "idempotencyKey" -]New value: +[ + "workspace" +]
- Added
open_run - Added
read_run_progress - Added
reauthenticate_cursor - Changed
reply_run7 fields changed- added
Input schema / properties / codexAllowedToolsAdded value: +{ + "description": "本次续接由 Codex 主进程决定的额外工具放行;省略时继承父运行,传空数组可撤销。", + "items": { + "enum": [ + "delete", + "task", + "mcp", + "generateImage" + ], + "type": "string" + }, + "maxItems": 4, + "type": "array" +} - added
Input schema / properties / targetLocationsAdded value: +{ + "description": "工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。", + "items": { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + "maxItems": 100, + "type": "array" +} - added
Input schema / properties / task / descriptionAdded value: +"任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。" - added
Input schema / properties / task / maxLengthAdded value: +4000 - added
Input schema / properties / timeoutMs / descriptionAdded value: +"Cursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。" - changed
Input schema / properties / timeoutMs / maximumPrevious value: -9007199254740991New value: +86400000 - changed
Input schema / properties / timeoutMs / minimumPrevious value: --9007199254740991New value: +1000
- Changed
start_run7 fields changed- added
Input schema / properties / codexAllowedToolsAdded value: +{ + "description": "仅由 Codex 主进程按任务范围决定的额外工具放行;高风险且超出用户既有授权时应先人工确认。", + "items": { + "enum": [ + "delete", + "task", + "mcp", + "generateImage" + ], + "type": "string" + }, + "maxItems": 4, + "type": "array" +} - added
Input schema / properties / targetLocationsAdded value: +{ + "description": "工作区内的文件、目录或行号位置列表,仅传位置不传内容;省略表示由 Cursor 按任务范围在工作区内定位。", + "items": { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + "maxItems": 100, + "type": "array" +} - added
Input schema / properties / task / descriptionAdded value: +"任务或审查范围与验收要求;禁止传源码正文、代码块或补丁,由 Cursor 在获授权工作区自行读取。" - added
Input schema / properties / task / maxLengthAdded value: +4000 - added
Input schema / properties / timeoutMs / descriptionAdded value: +"Cursor 任务总预算(毫秒);普通任务省略即可使用 24 小时默认值,硬上限同为 24 小时。" - changed
Input schema / properties / timeoutMs / maximumPrevious value: -9007199254740991New value: +86400000 - changed
Input schema / properties / timeoutMs / minimumPrevious value: --9007199254740991New value: +1000
- Added
view_run
10 tool updates
v0.1.0- First observed
authorize_workspace - First observed
cancel_run - First observed
doctor - First observed
get_run - First observed
list_models - First observed
list_runs - First observed
read_events - First observed
reply_run - First observed
start_run - First observed
wait_run
TDQS
Scored across 14 tools
Several tools target run observation and reading state (open_run, get_run, read_events, read_run_progress, view_run, wait_run), and their boundaries are subtle despite descriptions clarifying snapshot vs. streaming vs. live panel. Core run-control and auth tools are distinct, but the monitoring cluster invites misselection.
Almost all tools use snake_case verb_noun-style names such as list_models, start_run, and wait_run, with only doctor as a noun exception. The convention is predictable and readable.
14 tools is reasonable for auth, model discovery, workspace authorization, run lifecycle, and monitoring. The monitoring surface is somewhat heavy, but each tool maps to a defined operation.
The surface covers configuration/auth, model discovery, run creation, continuation, cancellation, listing, event reading, and progress monitoring. Explicit cleanup/prune/delete for persisted runs is missing, but the core lifecycle is otherwise covered.
Maintenance
Related MCP Connectors
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
Reliable async execution for agent tool calls: schema gating, retries, idempotency, audit trail.
Hosted TikTok ads MCP with OAuth, bounded reads, and prepare/confirm writes.
Scoped agent execution. Server-side credentials, policy, budgets and verifiable receipts.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables MCP clients to invoke Cursor SDK's agent runtime, run coding agents, list models, and continue conversations.41-
- AlicenseAqualityCmaintenanceEnables agents to submit and manage persistent, dependency-aware task graphs with immutable artifacts, resource reservations, durable event streaming, and retryable process execution over MCP.12MIT
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients to run declarative agents and DAG workflows as plain tools, with parallel nodes, review loops, and per-run least-privilege sandboxing.198 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables durable, least-privilege handoff of bounded tasks from a local producer to a configured MCP worker, with atomic claim, complete, and fail operations under expiring leases.MIT