dsh-relay
dsh-relay
Englisch | 中文
Web-Profil-Helfer, der den Codex-MCP-Startpfad diagnostiziert und die Host-Konfiguration schreibt, damit Codex, Cursor oder Claude Code dsh --profile codex starten können. Er führt den MCP- Stdio-Server nicht im Web-Prozess aus.
In ein Profil installieren:
dsh plugin --profile web add github:tonytanglab/deepseek-harness-relay-mcpoder installieren Sie das GitHub-Release-Tarball für ein vorgebautes lib/. Starten Sie dsh web nach der Installation neu. Führen Sie dann /relay-setup aus oder bitten Sie den Agenten, relay_doctor und relay_write_mcp_config aufzurufen.
Was es tut
apply registriert einen Slash-Befehl und zwei Tools im Web-Profil:
/relay-setupgibt die Doctor-Zusammenfassung und das MCP-Start-JSON aus. WennmcpConfigPathgesetzt ist, schreibt es auch dieses JSON. Zusätzliche Eingabe wird abgelehnt.relay_doctorprüft, obprocess.argv[1]ein absoluter existierender dsh-Eintrag ist, obprocess.execPathexistiert und ob der Credentials-Pfad existiert, ohne Credential-Inhalte zu lesen und ohne MCP stdio zu starten.relay_write_mcp_configerstellt den npx-Startblock aus dem Codex-Leitfaden. Es schreibt den Block, wennpatheine absolute JSON-Datei ist, und führtmcpServerszusammen, sodass andere Server erhalten bleiben.
Der generierte Befehl ist npx --yes --package=@deepseek-ai/dsh@0.1.0-rc.5 -- dsh --profile codex. Der Helfer verwendet niemals npx.cmd oder shell: true und importiert niemals StdioServerTransport.
Related MCP server: BoundedRelay
Konfiguration
Feld | Standard | Bedeutung |
|
| Schlüssel unter |
| unset | Absoluter JSON-Pfad, den |
|
| Leer liest |
|
| Gemeinsames Credentials-Dokument; auch |
|
| Codex-Service-Homes; auch |
|
| Festgelegtes Paket in generierten npx-Argumenten |
|
| Standard-Host für |
Benutzer überschreiben Zeilen im Profil cordis.patch.yml. Ungültige Konfiguration führt zu fehlgeschlagenem Plugin-Laden.
Nach der Einrichtung
Richten Sie den MCP-Host auf den gedruckten Startblock aus. Sitzungsstart, Warten, Steuern und Abbrechen bleiben bei dsh --profile codex, dokumentiert in @deepseek-ai/dsh-mcp-codex und dem Codex-Leitfaden. Das gebündelte Skill benennt diesen Workflow.
Modellerfahrung
Tool-Schema
Was das Modell sieht
relay_doctor hat ein leeres Parameterobjekt. relay_write_mcp_config erfordert host als codex, cursor oder claude-code und akzeptiert optionalen absoluten path.
Token-Effekt
Feste Schema-Kosten bei jeder Anfrage, bei der die Tools sichtbar sind.
KV-Cache-Effekt
Präfix-stabil, solange die Definitionen und die Sichtbarkeit unverändert sind.
Tool-Aufrufverlauf und Ergebnis
Was das Modell sieht
relay_doctor gibt JSON mit ok, Launcher direct/exists/shell: false, Workspace-Roots und Credentials path/exists zurück. Es enthält niemals Credential-Dateiinhalte. relay_write_mcp_config gibt { written, path, host, serverName, config } zurück, dessen config.args --profile und codex enthalten. Der Erfolgstext von /relay-setup beginnt mit DSH Relay is loaded in this Web profile. It does not run the MCP stdio server here. Zusätzliche Eingabe gibt The /relay-setup command does not accept extra input. zurück.
Token-Effekt
Pro Aufruf, begrenzt durch den JSON-Bericht und den Startblock.
KV-Cache-Effekt
Unabhängig von früheren Durchläufen.
Bekannte Einschränkungen und zurückgestellte Arbeiten
Kein prozessinterner MCP stdio — Die Installation dieses Bundles in das Web-Profil macht die elf Codex-MCP-Tools nicht verfügbar. Diese bleiben bei
dsh --profile codex.Keine Einstellungs-UI — Es gibt kein
dsh.client-Formular in v0.1.0;/relay-setupund die beiden Tools sind der Konfigurationspfad.
Available Tools
25 toolscancel_runCancel a Harness runCDestructiveIdempotent
Request cancellation through the public Host API.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true, readOnlyHint=false, and idempotentHint=true. The description adds the subtle behavioral trait that this is a 'request' for cancellation rather than an immediate cancel, but it does not explain what happens to the run, whether cancellation is reversible, or any downstream effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler or repetition. It is front-loaded and easy to parse, though it is arguably too brief to carry the behavioral and usage context needed for a destructive operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with no output schema and no usage guidance, the description is underspecified. It omits cancellation semantics, prerequisites, effects on the run, and any indication of what the agent should expect after invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning about runId or idempotencyKey. The schema itself defines formats and constraints, but the description does not explain the purpose of either parameter, when idempotencyKey should be used, or how runId relates to a cancellation flow.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Request cancellation') and a clear resource ('a Harness run'), which matches the title and tool name. It is distinguishable from siblings like 'start_run' and 'steer_run', though it does not explicitly differentiate itself from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no when-to-use or when-not-to-use guidance and does not mention alternatives. An agent is not told when cancellation is appropriate, what conditions are required (e.g., run must be active), or how this differs from related operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
doctorCheck Harness Relay MCP and HarnessARead-onlyIdempotent
Check the external Harness Host and relay policy without reading credentials.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is well covered. The description adds meaningful behavioral context beyond annotations, especially that it operates 'without reading credentials,' which is important for security-conscious agents. This extra disclosure makes the behavior more transparent than the schema alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence that front-loads the tool's purpose and includes the most critical security-related caveat ('without reading credentials'). Every word earns its place, and the title provides useful additional framing without bloating the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool, the description is mostly adequate and the annotations carry the safety context. However, there is no output schema and the description does not indicate what the return payload looks like, whether it is a simple status, a policy report, or a boolean pass/fail. This missing expected-output information leaves an agent slightly uncertain about how to interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and required parameters are zero, so the schema has 100% coverage trivially. There are no parameter semantics to document, and the description correctly avoids inventing meaningless parameter details. Baseline of 4 is appropriate for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Check') and identifies a clear resource ('external Harness Host and relay policy'). It also distinguishes itself from sibling setup_doctor by framing this as a verification action rather than a setup action. It loses one point because it does not state what a successful check returns or what exact conditions are being validated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for verifying external Harness connectivity and relay policy, and the sibling name setup_doctor suggests an alternative for installation/configuration. However, there is no explicit when-to-use versus when-not-to-use guidance, and no mention of when to prefer sibling tools such as status_run or get_operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_operationRead a relay operationARead-onlyIdempotent
Read the durable idempotent operation record for a start, reply, steer, or cancel request.
| Name | Required | Description | Default |
|---|---|---|---|
| operationId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the 'durable' qualifier and the request-scope, but does not disclose behavior like not-found handling or the exact response shape. It contributes some context beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. The verb and object appear first, the scoping qualifiers follow, and every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with readOnly and idempotent annotations and no output schema, the description is nearly complete: it names the resource and the exact request types it covers. It stops short of describing response contents, but 'Read the ... operation record' sufficiently implies the return value for this simple read operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented operationId parameter. It partially does by indicating that the ID refers to a durable operation record for start/reply/steer/cancel requests. However, it does not explain where the operationId comes from or otherwise enrich the parameter meaning much beyond the schema's name and uuid format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a specific resource ('durable idempotent operation record'), and narrows the scope by enumerating the request types: start, reply, steer, or cancel. This makes the tool easy to distinguish from run-focused siblings like get_run or status_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when you need the operation record for a start, reply, steer, or cancel request. It does not explicitly name alternatives or state when not to use it, but the request-type enumeration is strong contextual guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_runRead a Harness runARead-onlyIdempotent
Preferred tool to reconcile and return one run snapshot; replaces the deprecated status_run alias. If status is running or unknown, hostPollContract.hostMustCallWaitRunAgain is true and the host must not conclude.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description's job is to add context beyond those. It does so well by disclosing the polling contract: a non-terminal status forces hostMustCallWaitRunAgain=true and forbids conclusion. It also notes the deprecation of status_run, which helps an agent understand expected behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence establishes purpose and preference, and the second delivers the critical behavioral contract. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool, the description covers the essential safety and polling behavior. It lacks an explicit return-value description and does not explain runId, but the 'run snapshot' phrase plus the schema's strongly typed runId keep the definition mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate by explaining the runId parameter or how it is used. It never mentions runId, format expectations, or how the run snapshot is selected. The schema itself is self-descriptive with a UUID pattern, but the description adds no parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'reconcile and return one run snapshot' and notes it replaces the deprecated status_run alias. It identifies a specific resource (a Harness run) and a read operation, though it does not explicitly distinguish itself from siblings like get_run_summary or open_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It frames get_run as the 'Preferred tool' for this task and calls out the deprecated status_run alternative. It also gives a concrete post-call condition: if status is running or unknown, the host must not conclude and must call wait_run again. It lacks explicit exclusions for other sibling tools, but the context is fairly clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_summaryGet a structured Harness run summaryBRead-onlyIdempotent
Project the current durable run snapshot into stable status, model, permission, elapsed time, and next-action fields.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | No | |
| runId | Yes | |
| usage | No | |
| status | Yes | |
| progress | No | |
| provider | No | |
| attention | No | |
| elapsedMs | No | |
| startedAt | Yes | |
| updatedAt | Yes | |
| finishedAt | No | |
| nextAction | Yes | |
| permissionMode | Yes | |
| reasoningEffort | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotenceHint=true, so the safety profile is clear. The description adds the context that the summary is projected from a durable snapshot into stable fields, but it does not disclose edge behaviors such as errors, missing runs, or staleness. This is acceptable given the annotations but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It is compact and places the core action first. However, words like 'Project' and 'durable run snapshot' are slightly jargon-heavy, which reduces immediate readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a single parameter, generous annotations, and an output schema, so return values need not be explained. Still, the description leaves gaps around when to select this tool over closely related run tools, and it provides no context about how the summary relates to the live run state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention the runId parameter at all. The parameter name and UUID format are somewhat self-explanatory, but the description does not compensate for the total lack of schema-level description coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Project') and a specific resource ('current durable run snapshot'), and lists the stable output fields. It is clear about what the tool does, but it does not explicitly distinguish itself from the sibling get_run, which could also plausibly return run information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus get_run, status_run, open_run, or wait_run. The phrase 'durable run snapshot' hints at a read-oriented summary, but there is no explicit when-to-use, alternative, or exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_capabilitiesList Harness run optionsARead-onlyIdempotent
List native Provider/model/reasoning, agent preset, and permission preset options without guessing names.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior, so the bar for extra disclosure is lower. The description adds that the tool returns canonical option names rather than guesses, which is useful, but it does not describe output format, ordering, or any environment-specific behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and object, lists the option categories compactly, and closes with a meaningful behavioral qualifier. Every word adds value and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only discovery tool, the description is complete: it names exactly what the agent will learn and gives no ambiguous invocation requirements. No output schema is present, but the description adequately conveys the return concept — available option names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the input schema is intentionally empty and there is no parameter burden to document. Baseline for zero-parameter tools is 4; the description adds relevant semantic context by naming the categories of values the tool exposes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), resource (native Provider/model/reasoning, agent preset, and permission preset options), and a clear purpose ('without guessing names'). It distinguishes this tool from sibling list tools like list_services, list_workspaces, and list_runs by naming the exact option categories it covers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without guessing names' implies this tool should be used when the agent needs authoritative option values before configuring or running something. It provides clear context for when to call it, though it does not explicitly name alternative tools or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsList Harness runsCRead-onlyIdempotent
Reconcile and list runs restored from durable relay state.
| Name | Required | Description | Default |
|---|---|---|---|
| serviceId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this read-only, idempotent, and non-destructive, so the description need not restate that. It adds a data-source context ('restored from durable relay state') but does not clarify what 'reconcile' does or how it relates to the read-only guarantee.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-sentence description is efficiently front-loaded and contains no filler. However, its brevity comes at the cost of clarity around 'reconcile'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should indicate what a caller receives (a list of runs) and whether serviceId filters results; neither is stated. The tool's role relative to other run tools is also underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter serviceId is not mentioned in the description at all, and schema coverage is 0%. Since the schema only provides a UUID format, the description fails to explain the parameter's role (e.g., filtering runs by service) or its optionality.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title and description clearly identify a list operation for runs. The phrase 'reconcile and list' adds ambiguity, and the description doesn't explicitly distinguish this from siblings like get_run or status_run, so it stops short of a top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to choose list_runs over the many run-related siblings (get_run, get_run_summary, status_run, open_run, wait_run). There is also no mention of whether the optional serviceId should be used to scope the list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_servicesList Harness attachmentsBRead-onlyIdempotent
List Host attachments restored from durable relay state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide the safety profile (readOnly, idempotent, non-destructive), so the description only needs to add extra behavioral context. It adds an origin detail ('restored from durable relay state') but does not explain what restoration means for the returned data or what behaviors to expect. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It loses a point because the 'Host' vs 'Harness' terminology inconsistency and the obscure 'durable relay state' phrase reduce clarity despite the brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only list tool with rich annotations, the description is close to sufficient, but it leaves the relationship between 'services', 'Host attachments', and 'durable relay state' unexplained. Without an output schema or return-value details, an agent may still be unsure what exactly will be listed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100%, so the baseline is 4. The description adds no parameter-specific meaning because there are no parameters to describe, and none are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action and resource ('List Host attachments'), but the resource is ambiguous and inconsistent with the tool name ('list_services') and title ('List Harness attachments'). It is not a tautology, but the meaning of 'Host attachments restored from durable relay state' is vague for an agent deciding what this tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives like list_workspaces or list_capabilities. The description does not mention any conditions, exclusions, or sibling tools, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspacesList Harness workspacesARead-onlyIdempotent
List the native Harness workspace registry used by Relay routing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| mode | Yes | |
| roots | Yes | |
| workspaces | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds no additional behavioral details beyond the fact that it lists something, and it does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that leads with the action and resource, then adds the relevant routing context. It contains no redundant wording and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple zero-parameter list operation with no nested objects, a well-defined output schema, and annotations that cover idempotence and safety. The description is sufficient for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is nothing to document. A baseline score of 4 is appropriate because the description cannot add parameter meaning that does not exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a clear resource ('native Harness workspace registry'), and a context qualifier ('used by Relay routing'). It is clearly about listing workspaces, not workspace sessions or services, though it does not explicitly name a sibling for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'used by Relay routing' gives some contextual hint about when this tool is relevant, but the description does not explicitly state when to prefer this tool over alternatives like list_workspace_sessions or list_services. The usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspace_sessionsList sessions in a Harness workspaceARead-onlyIdempotent
List reusable native sessions accounted to one registered Harness workspace without reading conversation content.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| sessions | Yes | |
| workspace | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false), so the bar is lower. The description adds genuine value by disclosing that the tool deliberately avoids reading conversation content, which is a meaningful behavioral boundary not present in annotations. No contradiction found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tightly worded sentence front-loaded with the action verb. Every phrase earns its place: reusable, native, registered, and without reading conversation content each add a distinct qualifier with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only list tool with a rich annotation set and an output schema, the description is largely sufficient. The minor gaps are the absence of ordering or pagination notes and no explicit alternative routing, but these are mitigated by the output schema and the simplicity of the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaning beyond the bare string schema by specifying that the workspace parameter must reference a 'registered Harness workspace,' but it doesn't explain how to obtain or format the identifier. The compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (List), a precise resource (reusable native sessions), a scope (one registered Harness workspace), and a meaningful constraint (without reading conversation content). It is clearly distinct from sibling tools that list services, workspaces, runs, or capabilities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'one registered Harness workspace' qualifier implies the precondition that the workspace must already be registered, and the 'without reading conversation content' phrasing hints at a non-content-oriented use case. However, no explicit when-to-use or when-not-to-use guidance is given, and no sibling alternatives are named, so routing relies on inference from resource names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_runOpen a Harness runAIdempotent
Open the native Harness Web session in the operating system default browser. Call only when the user explicitly asks to open or show the Harness page.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It adds an observable external behavior beyond the structured hints: a browser session is opened in the OS default browser. It also communicates that this is user-visible, not a background/API-only action, and gives an explicit guardrail for when it should be invoked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description uses two short sentences with no filler: one sentence states the action and the other states the when-condition. Every clause contributes to correct tool invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter, side-effect tool, this is nearly complete: the agent knows the behavior, the trigger condition, and the schema supplies the required runId. It misses only the parameter-level explanation, which prevents a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain that the required runId is the run whose Harness Web page should be opened. Although the parameter name and schema format communicate part of the meaning, the description itself adds no semantic value for the one parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the precise behavior: opening a native Harness Web session in the operating system default browser. It distinguishes this from run execution, status, or service tools by naming the exact launch mechanism and target.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description sets an explicit trigger condition: only call when the user explicitly asks to open or show the Harness page. It does not name alternative sibling tools, but the gate is narrow enough for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_serviceOpen DeepSeek HarnessBIdempotent
Open the stable loopback Harness page for an attached service.
| Name | Required | Description | Default |
|---|---|---|---|
| serviceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already indicate this is not read-only and operates in an open world, yet the description does not disclose what state change or external effect occurs when the page is opened. It adds no behavioral context beyond the basic action name, and there is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler, redundant examples, or restatement of the schema. Every word contributes to identifying the target resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with one parameter and annotations, this is minimally viable: it names the action, target, and precondition. Clear gaps remain—'Harness page', 'stable loopback', and 'attached' are domain terms not explained, and no sibling differentiation is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides only the type and UUID format for serviceId, so the description's 'attached service' wording supplies minimal semantic meaning: the ID must refer to an attached service. It does not explain how to resolve the ID or what happens for invalid service IDs, but for a single self-named parameter this is marginally adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Open') and a specific resource ('the stable loopback Harness page') scoped to an attached service, making the tool's core function clear. It does not explicitly contrast with sibling open_run, but the service-vs-run target makes the distinction reasonably clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool over alternatives such as open_run, start_service, or list_services. The phrase 'attached service' implies a precondition but does not state it as a requirement or explain how to confirm attachment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_notificationsReplay retained DSH Relay notificationsARead-onlyIdempotent
Read the bounded in-process notification projection after a cursor; cursor gaps return explicit resynchronization metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| nextCursor | Yes | |
| notifications | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly and idempotent annotations, the description discloses meaningful behavior: the projection is bounded and in-process, and cursor gaps produce explicit resynchronization metadata. This adds substantive context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the action and resource, then adds the key behavioral nuance. Every clause contributes semantic value with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return shape, annotations cover safety, and the description explains the main cursor behavior. The only notable gap is the behavior when cursor is omitted, which is not addressed in the description or schema descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry cursor meaning. It does so by explaining that reads happen 'after a cursor' and that gaps trigger resynchronization metadata. It does not explain initial-cursor omission behavior, but the single optional parameter is otherwise well contextualized.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a specific resource ('bounded in-process notification projection'), and clearly scopes the operation with cursor semantics. This distinguishes it from the unrelated sibling tools and from any generic 'list notifications' operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: read notifications after a cursor, with specific resynchronization behavior on cursor gaps. It does not explicitly enumerate when-not-to-use or alternatives, but no close sibling alternative exists among the provided tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcile_operationReconcile a relay operationARead-onlyIdempotent
Compare an uncertain operation with durable Harness events without submitting a duplicate request.
| Name | Required | Description | Default |
|---|---|---|---|
| operationId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, and the description reinforces this by promising no duplicate request. It adds useful context by specifying reconciliation against 'durable Harness events', which tells the agent this is an event-based check rather than a live query.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single 14-word sentence that front-loads the verb and object, then packs the crucial side-effect guarantee at the end. There is no fluff, no repetition of the title, and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, annotation-rich tool, the description covers the core purpose, the source of truth, and the safety guarantee. It does not describe the return value or how to interpret the comparison result, and no output schema exists to fill that gap, so a small completeness gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter operationId is not described in the schema (0% coverage), and the description only refers to 'an uncertain operation' without explicitly mapping it to the required parameter. It adds the qualitative notion of uncertainty, but leaves the parameter semantics largely implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb 'Compare' and identifies the resource ('uncertain operation') and the reference data ('durable Harness events'). The phrase 'without submitting a duplicate request' distinguishes it from retry or creation tools, and it is clearly differentiated from siblings like get_operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'uncertain operation' signals the intended use case: when an operation's outcome is ambiguous and the agent needs to verify against durable events. It does not explicitly name alternatives such as get_operation or status_run, nor does it state when not to use it, so some inference remains.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcile_permissionsRestore a Harness session permission leaseBIdempotent
Retry restoration of the previous native permission preset for a session that requires attention.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint=true, destructiveHint=false, and readOnlyHint=false, so the safety profile is covered. The description adds the useful fact that it restores the previous native preset rather than setting arbitrary permissions, but it does not disclose side effects such as overwriting current custom permissions or any prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with the key action and object front-loaded, and it avoids repeating the title verbatim. 'That requires attention' is somewhat vague but does not add significant bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, idempotent restoration tool, the description plus annotations give a minimally workable picture. However, with no output schema, it does not say what a successful restoration returns, what signals that attention is required, or what state the session should be in before calling this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never explains sessionId, where to get it, or how it relates to the restore operation. The tool mentions 'session' generally, which weakly implies the parameter identifies the affected session, but the description does not compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retry restoration') and a specific resource ('previous native permission preset for a session'), making the tool's action clear. It does not explicitly differentiate from siblings like reconcile_operation or setup_doctor, but the permission-restoration purpose is identifiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for a session that requires attention' implies this is a corrective or retry operation, but it never explicitly states when to use it versus alternatives such as reconcile_operation or setup_doctor. There is no when-not-to-use guidance or mention of alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_runContinue a Harness sessionBDestructive
Submit a new queued turn to the completed run session, optionally selecting a different model, and track it as a new run. Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| model | No | ||
| runId | Yes | ||
| content | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| provider | No | ||
| writeScope | No | Paths Harness may modify. Must be empty for read-only runs; this is an instruction declaration in addition to native permissions. | |
| openBrowser | No | Keep false unless the user explicitly asks to open the Harness page in the OS browser. | |
| excludedPaths | No | Paths excluded from contextual reading. This is an instruction declaration, not filesystem enforcement. | |
| reviewTargets | No | Subjects to assess. This is not a read whitelist; supply together with contextReadScope. | |
| idempotencyKey | No | ||
| reasoningEffort | No | ||
| contextReadScope | No | Workspace paths Harness may search and read for evidence. Use only the targets when the user explicitly requests a target-only review. | |
| permissionPreset | No | ||
| authorizationBasis | No | Truthful evidence that the user explicitly selected Harness or this named Harness model to inspect the stated workspace scope. The current explicit request is sufficient authorization for that selected provider to process in-scope reads; do not ask the user to repeat authorization solely because provider processing is external. This field does not broaden the workspace, scope, permissions, destination, or allowed external actions. | |
| confirmedDangerousPermission | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses important behavioral traits beyond annotations: it warns that the selected Harness model may send content to its external provider, and it explicitly states that path fields are instructions not filesystem enforcement. These add significant context that the annotations (openWorldHint, destructiveHint) do not fully specify.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is verbose and repeats the same long paragraph verbatim in the schema for task, content.text, and the content array. This duplication is wasteful and hinders quick scanning, even though the main description is front-loaded with the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 15 parameters, no output schema, and a complex operational context, the description covers only a subset (mostly path and data-handling rules). It omits guidance on model selection, permission presets, idempotency, and dangerous confirmation, leaving the agent with significant gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 53%, and the description compensates for key path-based parameters (reviewTargets, contextReadScope, excludedPaths, writeScope) by explaining their semantics and constraints. However, it does not address other parameters like model, provider, reasoningEffort, permissionPreset, or idempotencyKey, so it is only partially compensating.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Submit a new queued turn to the completed run session' and clarifies it tracks as a new run, which distinguishes it from starting fresh runs. However, the phrase 'completed run session' is slightly ambiguous and does not explicitly name sibling tools like start_run or steer_run, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (continuing a completed run) but does not explicitly state when to use this tool versus alternatives such as start_run or steer_run, nor does it mention any exclusions or prerequisites. The focus is on parameter semantics rather than routing the agent to the correct tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_doctorDiagnose a Harness Relay MCP client setup planARead-onlyIdempotent
Return a machine-readable setup report from explicitly supplied probes without reading or modifying client configuration.
| Name | Required | Description | Default |
|---|---|---|---|
| facts | No | ||
| scope | Yes | ||
| client | Yes | ||
| platform | Yes | ||
| relayEntry | Yes | ||
| environment | No | ||
| homeDirectory | Yes | ||
| nodeExecutable | Yes | ||
| endpointDescriptor | Yes | ||
| workspaceDirectory | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| plan | Yes | |
| scope | Yes | |
| checks | Yes | |
| client | Yes | |
| status | Yes | |
| planReady | Yes | |
| schemaVersion | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds a useful behavioral guarantee: it does not read or modify client configuration. It also clarifies the operational mode, 'from explicitly supplied probes', and the output style, 'machine-readable setup report', providing context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly written sentence that fronts the core purpose and the highest-value constraint upfront. Every word earns its place; there is no fluff or repetition of schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex tool with ten parameters, seven required, and a nested facts object, yet the description gives no guidance on how to assemble or interpret those parameters. Even though output schema and annotations exist, an agent would likely struggle to call this tool correctly without understanding what 'probes' are and how the required strings relate to the diagnostic process.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters, but it only alludes to 'explicitly supplied probes' without detailing any of the ten parameters. It does not explain what 'facts' means, how the booleans should be set, or what values like relayEntry or endpointDescriptor represent, leaving the agent to infer entirely from names and enums.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb ('Return'), a concrete resource ('a machine-readable setup report'), and a distinctive scope condition ('from explicitly supplied probes without reading or modifying client configuration'). This clearly differentiates it from sibling diagnostic or planning tools like doctor or setup_plan by emphasizing that no client configuration is read or changed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when a non-invasive, probe-based setup report is needed and no configuration access is desired. However, it does not explicitly name an alternative, state when NOT to use this tool, or explain how it contrasts with sibling tools such as setup_plan or doctor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setup_planPlan Harness Relay MCP client setupBRead-onlyIdempotent
Generate a validated, no-write MCP configuration patch for Codex, Claude Code, Cursor, or OpenCode V2.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | Yes | ||
| client | Yes | ||
| platform | Yes | ||
| relayEntry | Yes | ||
| environment | No | ||
| homeDirectory | Yes | ||
| nodeExecutable | Yes | ||
| endpointDescriptor | Yes | ||
| workspaceDirectory | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| patch | No | |
| ready | Yes | |
| issues | Yes | |
| actions | Yes | |
| launcher | No | |
| detection | Yes | |
| writeAuthorized | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds 'no-write' and 'validated,' which are consistent with the annotations but do not provide rich behavioral detail such as how validation works or what the patch contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly written sentence with no filler. It front-loads the core action and deliverable while naming the supported targets, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, 7 required, 0% schema documentation, and many sibling tools, the description is too minimal. It explains what the tool produces but not where it fits in a setup workflow, what inputs are needed, or how to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% while there are 9 parameters, including 7 required ones. The description does not explain any parameter meaning, relationships, or expected values, so an agent cannot infer how to fill in fields like endpointDescriptor, relayEntry, or scope.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Generate'), a specific deliverable ('validated, no-write MCP configuration patch'), and the target clients (Codex, Claude Code, Cursor, OpenCode V2). This clearly distinguishes setup_plan from the sibling setup_doctor and other operational tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for planning a client setup without applying changes, but it gives no explicit when-to-use guidance, prerequisites, or alternatives. With sibling tools like setup_doctor and doctor present, an agent is not told how to choose between them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_reviewDispatch a read-only Harness reviewA
Create or reuse a native Harness session with the permission preset fixed to read-only. Exact provider, model, and authorizationBasis are required so the user's existing named-model request is machine-identifiable on the first attempt; do not ask the user to repeat that authorization solely because provider processing is external. Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. After start succeeds, share webUrl and keep wait_run until succeeded/failed/cancelled/needs_attention. The calling agent MUST read assistantText before claiming the review is done. If the parent user asked to review then fix, apply accepted findings only after the run is terminal; do not treat a still-running review as finished.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| model | Yes | Exact model returned by list_capabilities. Required so approval can identify the content-processing destination. | |
| content | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| provider | Yes | Exact provider returned by list_capabilities. Required so approval can identify the content-processing destination. | |
| sessionId | No | ||
| workspace | Yes | Authorized absolute Harness workspace root; Harness reads named in-scope files from here. | |
| writeScope | No | Paths Harness may modify. Must be empty for read-only runs; this is an instruction declaration in addition to native permissions. | |
| agentPreset | No | ||
| openBrowser | No | Keep false unless the user explicitly asks to open the Harness page in the OS browser. | |
| sessionMode | No | ||
| excludedPaths | No | Paths excluded from contextual reading. This is an instruction declaration, not filesystem enforcement. | |
| reviewTargets | No | Subjects to assess. This is not a read whitelist; supply together with contextReadScope. | |
| idempotencyKey | No | ||
| reasoningEffort | No | ||
| contextReadScope | No | Workspace paths Harness may search and read for evidence. Use only the targets when the user explicitly requests a target-only review. | |
| authorizationBasis | Yes | Truthful evidence that the user explicitly selected Harness or this named Harness model to inspect the stated workspace scope. The current explicit request is sufficient authorization for that selected provider to process in-scope reads; do not ask the user to repeat authorization solely because provider processing is external. This field does not broaden the workspace, scope, permissions, destination, or allowed external actions. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the sparse annotations by disclosing that the Harness model may send read content to an external provider, that fields are instructions rather than enforced filesystem isolation, and that source text must never be embedded. It also clarifies the post-run lifecycle and the requirement to read assistantText.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but purposeful, with important caveats about path-reference-only usage, external processing, and terminal-state handling. It is one long paragraph rather than structured bullets, but each sentence earns its place for a tool with this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter tool with no output schema and minimal annotations, the description is remarkably complete: it explains required fields, parameter semantics, data-flow caveats, success behavior, and post-run obligations. The caller knows what to provide, what will happen, and how to avoid prematurely reporting completion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds real meaning to key parameters: reviewTargets are not a read whitelist, contextReadScope defines where Harness may search, writeScope must be empty for read-only runs, and authorizationBasis should not require repeated user confirmation. However, optional parameters like sessionId, agentPreset, reasoningEffort, and sessionMode remain unexplained despite moderate schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title and first sentence clearly state a specific action: create or reuse a Harness session with a read-only permission preset to dispatch a review. This distinguishes it from sibling tools like start_run or steer_run because it names the review workflow and the read-only constraint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operating context: exact provider/model/authorizationBasis are required, task parameters are path-reference-only, and the caller must wait for terminal status and read assistantText before claiming completion. It does not explicitly name alternatives or say 'use this instead of X,' but the context is strong enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_runDispatch a Harness runADestructive
Create or reuse a native Harness session, select provider/model/reasoning, agent preset, and native permission preset, then submit the first task and return a stable session link. Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. Sharing webUrl is not completion: keep wait_run until a terminal status, then consume assistantText.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| model | No | ||
| content | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| provider | No | ||
| sessionId | No | ||
| workspace | Yes | Authorized absolute Harness workspace root; Harness reads named in-scope files from here. | |
| writeScope | No | Paths Harness may modify. Must be empty for read-only runs; this is an instruction declaration in addition to native permissions. | |
| agentPreset | No | ||
| openBrowser | No | Keep false unless the user explicitly asks to open the Harness page in the OS browser. | |
| sessionMode | No | ||
| excludedPaths | No | Paths excluded from contextual reading. This is an instruction declaration, not filesystem enforcement. | |
| reviewTargets | No | Subjects to assess. This is not a read whitelist; supply together with contextReadScope. | |
| idempotencyKey | No | ||
| reasoningEffort | No | ||
| contextReadScope | No | Workspace paths Harness may search and read for evidence. Use only the targets when the user explicitly requests a target-only review. | |
| permissionPreset | No | ||
| authorizationBasis | No | Truthful evidence that the user explicitly selected Harness or this named Harness model to inspect the stated workspace scope. The current explicit request is sufficient authorization for that selected provider to process in-scope reads; do not ask the user to repeat authorization solely because provider processing is external. This field does not broaden the workspace, scope, permissions, destination, or allowed external actions. | |
| confirmedDangerousPermission | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations, the definition discloses what would otherwise be invisible: task paths are instructions rather than enforced file-system isolation, Harness reads files from the workspace itself, and the selected model may send read content to its external provider despite Relay's loopback address. It also warns that webUrl sharing is not completion and flags the authorizationBasis requirement, which is strong behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the main purpose and then packs critical constraints into a dense but mostly purposeful paragraph. A few ideas reappear, such as path-reference-only and non-enforcement, and bullet formatting would improve scanability, but no major irrelevant content is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 18-parameter, potentially destructive tool with no output schema, the description covers lifecycle, security, egress, and permission semantics well. It omits concrete valid values for provider/model/reasoning/agentPreset and does not describe the returned run payload beyond a stable session link, so an agent may need list_capabilities or wait_run docs to fully proceed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema coverage, the description adds essential semantics for the highest-risk inputs: task must be path-reference-only, reviewTargets are not a read whitelist, contextReadScope defines where Harness may search, and writeScope must be empty for read-only runs. It does not explain every parameter such as provider, model, reasoningEffort, or agentPreset, leaving some vocabulary to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action—create or reuse a Harness session, submit the first task, and return a session link—so an agent understands what start_run accomplishes. It does not explicitly contrast start_run with close siblings such as start_review or reply_run, though the 'first task' wording implies the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear invocation context: creating vs reusing a session, needing a workspace and task instructions, and calling wait_run afterward instead of treating the returned webUrl as completion. It doesn't explicitly name alternatives like start_review or reply_run or state when not to use them, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_serviceAttach to DeepSeek HarnessAIdempotent
Attach an authorized workspace to the existing Harness Host. This never starts or modifies Harness.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | Yes | ||
| openBrowser | No | Keep false unless the user explicitly asks to open the Harness page in the OS browser. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, covering non-destructive behavior. The description adds the 'never starts or modifies Harness' clarification, which reinforces but does not significantly extend those hints. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action and a clear exclusion. Every sentence earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description captures core behavior and notes the prerequisite 'authorized workspace'. However, it does not mention return values, error conditions, or the impact of openBrowser beyond the schema description, leaving some operational gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: openBrowser has a description, but workspace does not. The tool description does not add any parameter-level detail (e.g., what constitutes a valid workspace identity). It fails to compensate for the missing workspace documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('attach') and resource ('authorized workspace to the existing Harness Host'), and explicitly clarifies it does not start or modify Harness. This distinguishes it from siblings like start_run and start_review without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to attach an authorized workspace to an existing host) and clarifies what it does not do (never starts or modifies Harness), but does not explicitly name alternative tools or provide conditions for choosing among related siblings. It relies on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
status_runDeprecated: get Harness run statusARead-onlyIdempotent
Deprecated compatibility alias; use get_run. Scheduled for removal in 0.3.0.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint true, and the description adds meaningful behavioral context by disclosing deprecation and the removal schedule. As a compatibility alias, failing to describe output details is acceptable since expected behavior is get_run's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short clauses with no wasted words; the deprecation and replacement are front-loaded. It earns every word it uses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple deprecated one-parameter read-only alias, the description provides essential routing and lifecycle information, and the schema fully constrains runId. It is adequate for an agent to invoke correctly, though output shape is left to get_run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides no parameter information, and schema description coverage is 0%. The single runId parameter is well-constrained by the schema's UUID format and pattern, reducing risk, but the description fails to compensate for the lack of schema descriptions as required at low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title clearly identifies the tool as 'get Harness run status' and the description explains it is a deprecated compatibility alias, so the action and resource are clear. It does not restate the behavior fully, but it differentiates itself from get_run by pointing to it as the replacement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs 'use get_run' and warns 'Scheduled for removal in 0.3.0', giving the agent a clear directive to choose the alternative over this deprecated tool. This is a precise when-not-to-use instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
steer_runSteer a Harness runADestructive
Durably insert a correction into an active run. Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| runId | Yes | ||
| content | No | Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address. | |
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral detail beyond the annotations: task parameters are path-reference-only, no source text may be embedded, Harness reads named files itself, fields are not enforced filesystem isolation, and the model may send read content to its configured provider despite Relay's loopback address. These are meaningful privacy and execution-model disclosures that a caller needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the essential purpose and then provides dense, valuable warnings. Each sentence adds a distinct constraint or behavioral fact; the only redundancy is that the same text appears verbatim in the schema's parameter descriptions, but the tool description itself is not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation on an active run, the description covers the critical safety and privacy aspects, including what must not be included in the payload and how Harness accesses files. It does not explain the return value or success/error behavior, but with no output schema and a clear purpose this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: runId and idempotencyKey have no schema descriptions, and the tool description does not compensate for those. However, the description does clarify the task/content parameter semantics thoroughly, explaining path-reference usage, reviewTargets, contextReadScope, and the prohibition on embedding source material, so it partially makes up for the gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause, 'Durably insert a correction into an active run,' uses a specific verb, resource, and state ('active run'), which clearly distinguishes it from siblings like start_run, reply_run, and cancel_run. It also names the core operation type ('correction') rather than merely restating the tool's name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it is for inserting a correction into an already active run, which implies it should not be used for starting, stopping, or merely observing runs. It does not explicitly name alternatives or exclusion conditions, so it misses the top score, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_serviceDetach from DeepSeek HarnessAIdempotent
Forget one relay attachment without stopping or changing the external Harness Host.
| Name | Required | Description | Default |
|---|---|---|---|
| serviceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover idempotency, non-readonly, and non-destructive intent. The description adds value beyond this by scoping the mutation precisely: exactly one relay attachment is affected, and the external Harness Host's state is preserved. This boundary disclosure is genuinely useful, though it stops short of describing the observable post-forget state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single 12-word sentence with zero waste: the action verb leads, followed by scope and the key exclusion. Every word earns its place and the critical 'does not stop the host' clarification is included without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is low-complexity (one parameter, no output schema), so the description covers the core semantics adequately. However, domain terms like 'relay attachment' and 'Harness Host' are unexplained, and the lifecycle relationship to siblings (how attachments are created via start_service/open_service) is missing, which could leave an agent unsure of prerequisites.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate for the single serviceId parameter. It partially does by tying the operation to 'one relay attachment,' implying serviceId selects the attachment. However, it never explicitly states that serviceId identifies which relay attachment to forget, leaving some ambiguity given the tool name says 'service' while the description says 'attachment.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('forget') and resource ('one relay attachment') and adds a clarifying boundary ('without stopping or changing the external Harness Host'), which usefully corrects the misleading 'stop_service' name. It implicitly distinguishes itself from stop-like operations, though it does not explicitly name a sibling to differentiate from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: use this when you want to detach/forget a single relay attachment while leaving the host untouched. However, no explicit when-to-use or when-not-to-use guidance is given, and no alternative tools are named (e.g., what to use if you actually need to stop the host or cancel an ongoing run).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_runWait for Harness progressARead-onlyIdempotent
Poll durable Host history for at most 30 seconds and return the latest snapshot plus hostPollContract. Call this native MCP tool directly; never poll through a temporary Node, PowerShell, Python, or shell client. A timeout is a slice, not completion. If hostPollContract.hostMustCallWaitRunAgain is true, you MUST call wait_run again immediately. Do not send a final user answer, mark the delegated task complete, or skip consuming assistantText while the run is still running. Unrelated shell or background-task notifications are not authorization to stop polling.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| timeoutMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description reveals that a timeout is only a slice not completion, that the caller must re-invoke wait_run when hostMustCallWaitRunAgain is true, and that unrelated notifications do not authorize stopping. This is exactly the kind of behavioral context annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries a distinct requirement or correction, from avoiding temporary poll clients to the prohibition on final answers during a running job. The key purpose is front-loaded, and the length is justified by the number of critical behavioral rules.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the polling protocol, return payload headline, and termination conditions well enough to call the tool correctly. It doesn's specify snapshot fields or hostPollContract structure, but no output schema is declared and the contract field is named, which is sufficient for a wait/loop tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no descriptions for runId or timeoutMs, so the description must compensate. It adds meaning to timeoutMs ('at most 30 seconds', 'timeout is a slice'), but says nothing about runId beyond the schema us type/format, leaving an agent to infer it identifies the run.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Poll') and resource ('durable Host history'), and names the deliverable ('latest snapshot plus hostPollContract'). This clearly differentiates wait_run from siblings like get_run or status_run, which are single-shot reads without a polling contract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs agents to call this native MCP tool directly rather than writing temporary poll clients, and it gives concrete loop rules: if hostMustCallWaitRunAgain is true, call again immediately, and do not finalize while the run is active. It doesn't name sibling alternatives, but the polling context and stop conditions provide enough usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.2.17- Changed
reply_run1 field changed- changed
Input schema / properties / authorizationBasis / descriptionPrevious value: -"Truthful evidence that the user explicitly selected Harness to inspect this workspace. Set only when the current conversation contains that instruction. This records evidence and does not expand permissions or guarantee approval."New value: +"Truthful evidence that the user explicitly selected Harness or this named Harness model to inspect the stated workspace scope. The current explicit request is sufficient authorization for that selected provider to process in-scope reads; do not ask the user to repeat authorization solely because provider processing is external. This field does not broaden the workspace, scope, permissions, destination, or allowed external actions."
- Changed
start_review4 fields changed- changed
Input schema / properties / authorizationBasis / descriptionPrevious value: -"Truthful evidence that the user explicitly selected Harness to inspect this workspace. Set only when the current conversation contains that instruction. This records evidence and does not expand permissions or guarantee approval."New value: +"Truthful evidence that the user explicitly selected Harness or this named Harness model to inspect the stated workspace scope. The current explicit request is sufficient authorization for that selected provider to process in-scope reads; do not ask the user to repeat authorization solely because provider processing is external. This field does not broaden the workspace, scope, permissions, destination, or allowed external actions." - added
Input schema / properties / model / descriptionAdded value: +"Exact model returned by list_capabilities. Required so approval can identify the content-processing destination." - added
Input schema / properties / provider / descriptionAdded value: +"Exact provider returned by list_capabilities. Required so approval can identify the content-processing destination." - changed
Input schema / requiredPrevious value: -[ - "workspace" -]New value: +[ + "workspace", + "provider", + "model", + "authorizationBasis" +]
- Changed
start_run1 field changed- changed
Input schema / properties / authorizationBasis / descriptionPrevious value: -"Truthful evidence that the user explicitly selected Harness to inspect this workspace. Set only when the current conversation contains that instruction. This records evidence and does not expand permissions or guarantee approval."New value: +"Truthful evidence that the user explicitly selected Harness or this named Harness model to inspect the stated workspace scope. The current explicit request is sufficient authorization for that selected provider to process in-scope reads; do not ask the user to repeat authorization solely because provider processing is external. This field does not broaden the workspace, scope, permissions, destination, or allowed external actions."
5 tool updates
v0.2.13- Changed
reply_run9 fields changed- added
Input schema / properties / authorizationBasisAdded value: +{ + "const": "explicit-user-request", + "description": "Truthful evidence that the user explicitly selected Harness to inspect this workspace. Set only when the current conversation contains that instruction. This records evidence and does not expand permissions or guarantee approval.", + "type": "string" +} - changed
Input schema / properties / content / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - added
Input schema / properties / contextReadScopeAdded value: +{ + "description": "Workspace paths Harness may search and read for evidence. Use only the targets when the user explicitly requests a target-only review.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" +} - added
Input schema / properties / excludedPathsAdded value: +{ + "description": "Paths excluded from contextual reading. This is an instruction declaration, not filesystem enforcement.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / openBrowser / descriptionAdded value: +"Keep false unless the user explicitly asks to open the Harness page in the OS browser." - added
Input schema / properties / reviewTargetsAdded value: +{ + "description": "Subjects to assess. This is not a read whitelist; supply together with contextReadScope.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" +} - changed
Input schema / properties / task / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address." - added
Input schema / properties / writeScopeAdded value: +{ + "description": "Paths Harness may modify. Must be empty for read-only runs; this is an instruction declaration in addition to native permissions.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "type": "array" +}
- Changed
start_review9 fields changed- added
Input schema / properties / authorizationBasisAdded value: +{ + "const": "explicit-user-request", + "description": "Truthful evidence that the user explicitly selected Harness to inspect this workspace. Set only when the current conversation contains that instruction. This records evidence and does not expand permissions or guarantee approval.", + "type": "string" +} - changed
Input schema / properties / content / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - added
Input schema / properties / contextReadScopeAdded value: +{ + "description": "Workspace paths Harness may search and read for evidence. Use only the targets when the user explicitly requests a target-only review.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" +} - added
Input schema / properties / excludedPathsAdded value: +{ + "description": "Paths excluded from contextual reading. This is an instruction declaration, not filesystem enforcement.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / openBrowser / descriptionAdded value: +"Keep false unless the user explicitly asks to open the Harness page in the OS browser." - added
Input schema / properties / reviewTargetsAdded value: +{ + "description": "Subjects to assess. This is not a read whitelist; supply together with contextReadScope.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" +} - changed
Input schema / properties / task / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address." - added
Input schema / properties / writeScopeAdded value: +{ + "description": "Paths Harness may modify. Must be empty for read-only runs; this is an instruction declaration in addition to native permissions.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "type": "array" +}
- Changed
start_run9 fields changed- added
Input schema / properties / authorizationBasisAdded value: +{ + "const": "explicit-user-request", + "description": "Truthful evidence that the user explicitly selected Harness to inspect this workspace. Set only when the current conversation contains that instruction. This records evidence and does not expand permissions or guarantee approval.", + "type": "string" +} - changed
Input schema / properties / content / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - added
Input schema / properties / contextReadScopeAdded value: +{ + "description": "Workspace paths Harness may search and read for evidence. Use only the targets when the user explicitly requests a target-only review.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" +} - added
Input schema / properties / excludedPathsAdded value: +{ + "description": "Paths excluded from contextual reading. This is an instruction declaration, not filesystem enforcement.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / openBrowser / descriptionAdded value: +"Keep false unless the user explicitly asks to open the Harness page in the OS browser." - added
Input schema / properties / reviewTargetsAdded value: +{ + "description": "Subjects to assess. This is not a read whitelist; supply together with contextReadScope.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "minItems": 1, + "type": "array" +} - changed
Input schema / properties / task / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address." - added
Input schema / properties / writeScopeAdded value: +{ + "description": "Paths Harness may modify. Must be empty for read-only runs; this is an instruction declaration in addition to native permissions.", + "items": { + "description": "Workspace-relative path, or an absolute path contained by the authorized workspace.", + "minLength": 1, + "type": "string" + }, + "type": "array" +}
- Changed
start_service1 field changed- added
Input schema / properties / openBrowser / descriptionAdded value: +"Keep false unless the user explicitly asks to open the Harness page in the OS browser."
- Changed
steer_run3 fields changed- changed
Input schema / properties / content / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - changed
Input schema / properties / task / descriptionPrevious value: -"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."New value: +"Task parameters are path-reference-only: provide the authorized workspace, reviewTargets, contextReadScope, excludedPaths, writeScope, and acceptance criteria. reviewTargets identify what to assess; they are not a read whitelist. contextReadScope declares where Harness may search and read supporting implementation, tests, configuration, and architecture material. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. These fields are task instructions, not enforced per-path filesystem isolation. The selected Harness model may send content it reads to its configured model provider even though Relay itself uses a loopback address."
4 tool updates
v0.2.9- Changed
reply_run3 fields changed- added
Input schema / properties / content / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - added
Input schema / properties / task / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."
- Changed
start_review4 fields changed- added
Input schema / properties / content / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - added
Input schema / properties / task / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions." - added
Input schema / properties / workspace / descriptionAdded value: +"Authorized absolute Harness workspace root; Harness reads named in-scope files from here."
- Changed
start_run4 fields changed- added
Input schema / properties / content / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - added
Input schema / properties / task / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions." - added
Input schema / properties / workspace / descriptionAdded value: +"Authorized absolute Harness workspace root; Harness reads named in-scope files from here."
- Changed
steer_run3 fields changed- added
Input schema / properties / content / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions." - changed
Input schema / properties / content / items / oneOfPrevious value: -[ - { - "properties": { - "text": { - "maxLength": 100000, - "type": "string" - }, - "type": { - "const": "text", - "type": "string" - } - }, - "required": [ - "type", - "text" - ], - "type": "object" - }, - { - "properties": { - "data": { - "minLength": 1, - "type": "string" - }, - "mediaType": { - "enum": [ - "image/png", - "image/jpeg", - "image/webp", - "image/gif" - ], - "type": "string" - }, - "name": { - "type": "string" - }, - "type": { - "const": "image", - "type": "string" - } - }, - "required": [ - "type", - "mediaType", - "data" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "text": { + "description": "Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions.", + "maxLength": 100000, + "type": "string" + }, + "type": { + "const": "text", + "type": "string" + } + }, + "required": [ + "type", + "text" + ], + "type": "object" + }, + { + "properties": { + "data": { + "minLength": 1, + "type": "string" + }, + "mediaType": { + "enum": [ + "image/png", + "image/jpeg", + "image/webp", + "image/gif" + ], + "type": "string" + }, + "name": { + "type": "string" + }, + "type": { + "const": "image", + "type": "string" + } + }, + "required": [ + "type", + "mediaType", + "data" + ], + "type": "object" + } +] - added
Input schema / properties / task / descriptionAdded value: +"Task parameters are path-reference-only: provide the authorized workspace, file or directory locations, scope, and acceptance criteria. Never embed source text, diffs, file dumps, encoded source, or repository archives. Harness reads named files from the authorized workspace itself. This rule is identical for read-only and write-capable permissions."
25 tool updates
v0.2.6- First observed
cancel_run - First observed
doctor - First observed
get_operation - First observed
get_run - First observed
get_run_summary - First observed
list_capabilities - First observed
list_runs - First observed
list_services - First observed
list_workspace_sessions - First observed
list_workspaces - First observed
open_run - First observed
open_service - First observed
read_notifications - First observed
reconcile_operation - First observed
reconcile_permissions - First observed
reply_run - First observed
setup_doctor - First observed
setup_plan - First observed
start_review - First observed
start_run - First observed
start_service - First observed
status_run - First observed
steer_run - First observed
stop_service - First observed
wait_run
TDQS
Scored across 25 tools
Most tools target distinct resources and actions (runs vs services vs workspaces), and the descriptions carefully differentiate get_run, wait_run, get_run_summary, and list_runs. The main ambiguity is the intentional status_run alias for get_run, plus start_run/start_review share similar task-parameter language, though their permission/preset distinction is clear.
The overwhelming majority follow snake_case verb_noun names like list_runs, start_service, cancel_run, and wait_run. The exceptions are bare doctor and setup_doctor, which break the verb_noun pattern but do not introduce mixed casing styles.
25 tools is at the high end of the borderline range, and many are needed for the run/relay lifecycle. Still, a deprecated alias (status_run), several setup/reconcile helpers, and notification/operation readers make the surface feel heavier than necessary.
The run lifecycle is well covered: create, review, reply, steer, get, wait, cancel, list, and reconcile. The gaps are minor—there is no explicit way to delete/forget a workspace session or run, and workspace management is read-only—but the core workflows have no dead ends.
Maintenance
Related MCP Connectors
Provision an OpenCode workbench and MCP stack on any Linux box, local or over SSH.
Enable secure connectivity between Sentry issues and debugging data, and LLM clients, using a Model Context Protocol (MCP) server.
Network, domain and website diagnostics for AI clients via MCP.
Collision-first MCP diagnostics for MCP, API, webhook and agent failures. Returns structured evidence, bounded repair guidance, agent compatibility checks and safe recovery routing.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceLaunches the official Codex CLI as a persistent MCP server so DeepSeek Harness can invoke Codex models without storing or transmitting ChatGPT credentials to third parties.MIT
- AlicenseNot gradedqualityBmaintenanceEnables Claude Code to delegate bounded, read-only analysis and isolated patch proposals to the locally installed Codex CLI over MCP, with sanitized live status, revision-aware polling, and reviewable results.1MIT
- AlicenseNot gradedqualityAmaintenanceEnables ChatGPT Web or another MCP client to inspect and develop a local workspace, read local Codex task history, and use the complete installed XcodeBuildMCP catalog through a small set of stable tools.MIT
- AlicenseAqualityCmaintenanceDelegates coding tasks to the locally installed DeepSeek Harness CLI, enabling MCP clients to run prompts through dsh and receive results with exit code and timing.1MIT