depot
Server Details
Inspect Depot builds, CI runs, job logs and Actions runners; retry or cancel CI runs.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
- Repository
- m190/usefulapi-mcp
- GitHub Stars
- 0
TDQS
Scored across 22 tools
Tools are partitioned by domain (CI, builds/projects/registry, GitHub Actions, usage) and each has a distinct resource+action. Overlaps like retry_job vs retry_failed_jobs or list_runs vs list_workflows are clearly differentiated by their descriptions, and navigation helpers like get_run_status explicitly tie the set together.
All tool names follow a consistent snake_case verb_noun pattern with a uniform `depot_` prefix. Service prefixes (`ci_`, `gha_`) are used predictably for those domains, while other tools omit a service prefix, yielding a coherent overall convention.
22 tools is on the high side but justified by the multi-service scope (CI, container builds, registry, GitHub Actions, usage). Each tool covers a distinct operation, so the set is slightly heavy but not bloated.
The surface is read-heavy and lacks update/delete operations for projects and registry images, and has no create for CI runs. These are notable lifecycle gaps that could block agents needing to modify or clean up resources, though core read and retry workflows are covered.
Available Tools
22 toolsdepot_ci_cancel_runCancel a CI runADestructiveInspect
WRITE — stops work in progress. Cancel a queued or running Depot CI run and all of its unfinished workflows, jobs and attempts. Cancelled jobs can be retried afterwards with depot_ci_retry_failed_jobs. Depot: CIService/CancelRun.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | CI run id, from depot_ci_list_runs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds real value beyond that: it enumerates what gets destroyed (workflows, jobs, attempts) and confirms the action is recoverable via retry. It does not say whether cancelling a finished run errors or is a no-op, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the mutation warning front-loaded, then the scope, then the recovery pointer. No filler and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema, the description covers safety intent, destruction scope, and the follow-up recovery tool. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single run_id parameter is already documented in the schema, including its provenance ('from depot_ci_list_runs'). The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (cancel) and resource (Depot CI run) plus the exact blast radius: 'all of its unfinished workflows, jobs and attempts.' This lets an agent distinguish it from read-oriented siblings like depot_ci_get_run_status without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'WRITE — stops work in progress' lead-in sets the context for when this tool applies, and it routes the recovery path to a named sibling (depot_ci_retry_failed_jobs). It stops short of explicit when-not guidance, such as behavior on an already-completed run.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_get_failure_diagnosisDiagnose a CI failureARead-onlyInspect
Get Depot's failure diagnosis for a CI run, workflow, job or attempt. For runs and workflows, identical failures are grouped by error fingerprint; for a job or attempt it is a focused diagnosis with evidence. Empty when no failure evidence is found. The best first call when a run failed. Depot: CIService/GetFailureDiagnosis.
| Name | Required | Description | Default |
|---|---|---|---|
| target_id | Yes | The run, workflow, job or attempt id. | |
| target_type | Yes | What target_id identifies. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already covers the safety profile, but the description adds real behavioral context beyond annotations: identical failures are grouped by error fingerprint for runs/workflows, job/attempt targets yield a focused diagnosis with evidence, and the result is empty when no failure evidence exists. That empty-result disclosure is genuinely useful for agent decision-making.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with purpose and containing no filler; the routing hint is placed last where it lands well. The trailing 'Depot: CIService/GetFailureDiagnosis.' identifier adds little for an agent, a minor tax on an otherwise efficient description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so qualitatively — describing grouped fingerprints, focused evidence, and the empty case. It could say more about the shape or size of the diagnosis, but nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning by explaining how behavior differs per target_type (grouped vs. focused diagnosis), which helps an agent choose between run, workflow, job and attempt. It stops short of clarifying the id format for each target type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get Depot's failure diagnosis for a CI run, workflow, job or attempt') and enumerates the exact target types. It also distinguishes itself from log/status siblings by describing what the diagnosis actually contains (fingerprint grouping vs. focused evidence).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly declares 'The best first call when a run failed', giving a clear triggering condition. It does not name competing alternatives (e.g. get_job_logs) or state exclusions, so it falls short of a full when/when-not routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_get_jobGet a CI jobARead-onlyInspect
Get one Depot CI job: status, conclusion, error message, timestamps, runner config (labels, image, CPUs, memory), matrix values, parent run/workflow context, dependencies and attempt history. Depot: CIService/GetJob.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | CI job id, from depot_ci_get_run_status. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, so safety is covered. The description adds real value beyond that by disclosing the breadth of returned data — error messages, timestamps, runner config details, matrix values, dependencies and attempt history — which matters because no output schema exists to reveal the payload shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence plus a provenance tag ('Depot: CIService/GetJob'). The parenthetical field enumeration is long but every item is informative rather than filler, and nothing is padded with generic prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-parameter tool with no output schema, the description is nearly complete: it enumerates the return surface that would otherwise be invisible to the agent. The missing piece is any indication of when to prefer this over its log/diagnosis siblings, but that is a routing gap rather than an incompleteness of the tool's contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and the schema documents it fully (100% coverage) including its origin in depot_ci_get_run_status. The description adds nothing about job_id, so the baseline of 3 applies; the schema carries the entire burden here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get one Depot CI job') and then enumerates exactly what the response contains (status, conclusion, error, runner config, matrix, dependencies, attempt history). This field list implicitly separates it from depot_ci_get_job_logs and depot_ci_get_failure_diagnosis, which return different payloads, but no sibling is named outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or alternative routing in the description. The only routing hint ('from depot_ci_get_run_status') lives in the schema's job_id description rather than the tool description, so the description itself offers no usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_get_job_logsGet a CI job's logsARead-onlyInspect
Get persisted log lines for a Depot CI job attempt, oldest first (step key/name, timestamp, line number, body). Pass job_id to read the job's latest attempt, or attempt_id for a specific one — exactly one. Pass nextPageToken back as page_token for more lines. Depot: CIService/GetJobAttemptLogs.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | CI job id — resolves to its latest attempt. | |
| attempt_id | No | A specific attempt id. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds real value beyond that: logs are 'persisted', returned 'oldest first', and pagination flows through nextPageToken/page_token, which an agent needs to avoid mis-reading the result set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then routing, then pagination in three tight sentences with no filler. The trailing 'Depot: CIService/GetJobAttemptLogs' identifier is mild overhead but arguably useful for tracing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned line fields (step key/name, timestamp, line number, body) and explaining the pagination cursor. Complete enough for correct invocation, though it omits any statement about maximum page size or truncation behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by stating the mutual-exclusion rule ('exactly one' of job_id/attempt_id) and by tying nextPageToken to the page_token parameter, which the schema does not express.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get persisted log lines for a Depot CI job attempt') plus ordering and returned fields. An agent can distinguish it from log-adjacent siblings like depot_gha_search_logs and depot_get_build_step_logs without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear selection guidance for the two lookup modes ('Pass job_id to read the job's latest attempt, or attempt_id for a specific one — exactly one') and tells the agent to pass nextPageToken back as page_token. It does not contrast the tool against sibling log tools, but the internal routing guidance is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_get_run_statusGet a CI run's status treeARead-onlyInspect
Get a Depot CI run's status with its nested workflows, jobs and attempts (ids, keys, statuses, error messages). This is how to find the workflowId / jobId / attemptId for the log, diagnosis and retry tools. Depot: CIService/GetRunStatus.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | CI run id, from depot_ci_list_runs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the safety profile is covered, and the description adds genuinely useful behavior beyond that: it discloses the exact return structure (nested workflows/jobs/attempts including ids, keys, statuses and error messages). With no output schema, this is the only place that shape is documented; only permission or rate-limit context is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, both earning their place: the first states what is returned, the second states what it is used for. The most decision-relevant content (the resource and its nested contents) is front-loaded with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description covers purpose, return shape, and downstream usage routing. The counterpart API name (CIService/GetRunStatus) is also included, leaving nothing an agent needs in order to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single run_id parameter is already documented there as 'CI run id, from depot_ci_list_runs.' The description adds no format, prefix or validation detail beyond the schema, so the baseline 3 for schema-driven parameter documentation applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Get) and resource (a Depot CI run's status) and enumerates the nested shape it returns (workflows, jobs, attempts with ids, keys, statuses, error messages). This distinguishes it from siblings like depot_ci_get_job or depot_ci_list_runs by making clear it is the hierarchical status tree, not a flat list or single entity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the purpose: 'This is how to find the workflowId / jobId / attemptId for the log, diagnosis and retry tools,' routing the agent to the correct downstream calls. It stops short of stating when NOT to use it or naming a competing sibling for the same data, so it is clear context without full exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_list_runsList Depot CI runsARead-onlyInspect
List Depot CI runs, newest first. Filter by status (default: running + queued — pass ["failed"] to find failures), repo, commit SHA prefix, trigger event, or pull request number (requires repo). Depot: CIService/ListRuns.
| Name | Required | Description | Default |
|---|---|---|---|
| pr | No | Pull request number. Requires repo. | |
| sha | No | Commit SHA prefix (1-40 hex chars). | |
| repo | No | Repository in owner/name format, e.g. depot/cli. | |
| status | No | Run statuses to include (ORed). Default on Depot's side: running and queued. | |
| trigger | No | Trigger event, e.g. push, pull_request, schedule, workflow_dispatch. | |
| page_size | No | Page size, 1-100. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true. The description adds valuable non-obvious behavior: results are ordered newest first and the default status filter is running+queued, which materially changes what an unqualified call returns. No pagination or return-shape disclosure, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph that front-loads what the tool does, then the filters, then the significant default behavior. No filler sentences; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param list tool with no output schema and readOnly annotations, the description covers purpose, filters, defaults, and a prerequisite. It omits any mention of pagination behavior, which the description could tie to page_size/page_token, leaving a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still earns credit by adding the actionable tip that ['failed'] finds failures and the repo-prerequisite for pr, which go beyond the schema text, though most of the status-default detail is repeated from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (Depot CI runs) plus ordering (newest first), clearly distinguishing it from sibling verbs like get_run_status or get_job. It does not, however, name the sibling alternatives it is distinct from, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides real selection context: the default status filter is running+queued, and passing ['failed'] surfaces failures; it also notes that pr requires repo. There is no explicit when-not-to-use guidance or direct reference to sibling tools, so it is strong but not complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_list_workflowsList Depot CI workflowsARead-onlyInspect
List Depot CI workflows across runs, newest first, with per-status job counts. Filter by name/path text, repo, status, trigger, SHA prefix or PR number. Depot: CIService/ListWorkflows.
| Name | Required | Description | Default |
|---|---|---|---|
| pr | No | Pull request number. | |
| sha | No | Commit SHA prefix (1-40 hex chars). | |
| name | No | Text to match in the workflow name or file path. | |
| repo | No | Repository in owner/name format. | |
| status | No | Workflow statuses to include (ORed). | |
| trigger | No | Trigger event, e.g. push or pull_request. | |
| page_size | No | Page size, 1-200. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds real behavior (results are ordered newest-first and enriched with per-status job counts), but says nothing about pagination despite exposing page_size, nor how many records a default call returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, with the core operation and ordering front-loaded and filters following. The trailing 'Depot: CIService/ListWorkflows' is internal implementation noise that consumes space without helping an agent invoke the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description usefully pre-announces the shape of results (ordered, counted by status). However, for a 7-parameter paged listing tool it omits pagination/limit behavior and default ordering scope, leaving gaps an agent must probe at runtime.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description loosely maps the filters ('name/path text', 'SHA prefix', 'PR number') but adds no semantics beyond the schema, and omits page_size entirely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('List Depot CI workflows') with scope ('across runs'), ordering ('newest first'), and an enrichment detail ('per-status job counts'). The agent can distinguish this from the sibling depot_ci_list_runs, which lacks the workflow-level grouping and counts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It enumerates the filter dimensions available, which implies when the tool is useful (narrowing a workflow search), but never states when to prefer it over siblings like depot_ci_list_runs or depot_ci_get_run_status, and offers no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_retry_failed_jobsRetry a workflow's failed jobsADestructiveInspect
WRITE — starts compute. Retry only the failed and cancelled jobs in a Depot CI workflow, plus skipped jobs that depend on them. Jobs that succeeded are left alone. Depot: CIService/RetryFailedJobs.
| Name | Required | Description | Default |
|---|---|---|---|
| workflow_id | Yes | Workflow id, from depot_ci_get_run_status or depot_ci_list_workflows. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint=true annotation, the description adds that this starts compute (a cost/resource implication) and delineates exactly which jobs are affected versus left alone, which is real behavioral value. It does not mention permissions, idempotency, or what happens to already-queued jobs, so it isn't exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences with the write/compute warning front-loaded, then the exact scope of mutation. Every clause carries information; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation tool with annotations covering the safety profile, the description covers what changes and what is left alone. With no output schema, a brief note on the return value or auth requirements would make it fully self-sufficient, but an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and schema description coverage is 100%, with the schema itself telling the agent to obtain workflow_id from depot_ci_get_run_status or depot_ci_list_workflows. The description adds no parameter meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retry) and resource (a workflow's failed jobs), and scopes it precisely: failed and cancelled jobs plus dependent skipped jobs, while succeeded jobs are untouched. This distinguishes it from the singular sibling depot_ci_retry_job without the agent needing to open a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the selection condition clear (use when you want to re-run only the failed/cancelled subset, not the whole workflow), which implicitly separates it from a single-job retry. It stops short of naming the alternative tool or stating when-not to use it, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_ci_retry_jobRetry a CI jobADestructiveInspect
WRITE — starts compute. Retry one failed or cancelled Depot CI job, plus any skipped jobs that depend on it. Fails if a dependent job already started; use depot_ci_retry_failed_jobs then. Depot: CIService/RetryJob.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job id to retry. | |
| workflow_id | Yes | Workflow id containing the job, from depot_ci_get_run_status. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description adds real behavior beyond that: 'WRITE — starts compute', the cascading retry of dependent skipped jobs, and the failure precondition when a dependent job already started. It doesn't cover auth or return semantics, but for a mutation tool this discloses the important side effects and pitfalls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the WRITE warning, then states action, cascade behavior, and the fallback alternative in three tight sentences with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and full parameter coverage, the description supplies the essentials: mutation nature, cascade scope, and the failure/alternative path. Return-value behavior is the only omission, which is minor given no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (job_id, workflow_id) are already documented, including the workflow_id provenance hint. The description adds no additional parameter semantics beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource: retry one failed or cancelled Depot CI job. It also specifies the cascade scope ('plus any skipped jobs that depend on it'), which distinguishes it from sibling retry/cancel tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative depot_ci_retry_failed_jobs and the exact condition that selects it ('Fails if a dependent job already started; use ... then'). This is explicit when-to-use and when-to-switch guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_create_projectCreate a projectADestructiveInspect
WRITE. Create a Depot container-build project (an isolated build cache). Idempotent by name: if an active project with the same name exists, that project is returned. Defaults: cache 50 GB / 14 days, hardware HARDWARE_16X32 on AWS builders. Depot: ProjectService/CreateProject.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Display name for the project. | |
| hardware | No | Builder size, e.g. HARDWARE_16X32 (16 CPU / 32 GB). | |
| region_id | Yes | Depot region or cloud connection, e.g. us-east-1 or eu-central-1. | |
| cache_keep_gb | No | Max cache size per architecture, GB (25-1000). Default 50. | |
| cache_keep_days | No | Days to keep cache entries since last use (1-30). Default 14. | |
| organization_id | No | Owning organization. Only needed for user tokens that belong to more than one org. | |
| build_timeout_minutes | No | Max duration of each build, minutes. Omit for no timeout. | |
| max_concurrent_builds | No | Builds per builder before autoscaling adds one. Omit to disable autoscaling. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only destructiveHint/title in annotations, the description carries real weight and delivers: it labels the operation as WRITE, discloses name-based idempotency and the returned-project behavior on collision, and lists the default cache/hardware configuration. It does not mention auth/org-token requirements or side effects like builder provisioning, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operation type ('WRITE.') followed by purpose, then idempotency semantics, then defaults — four tight clauses, each earning its place, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, 2-required creation tool with no output schema, the description covers purpose, write semantics, idempotency and defaults adequately; only the return shape and permission/org-token prerequisites are left implicit, which the schema partially covers via organization_id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description's default recap (50 GB / 14 days / HARDWARE_16X32) largely restates what the schema already documents per-parameter and adds no format or interaction detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a Depot container-build project') plus a clarifying gloss ('an isolated build cache'), which cleanly separates it from depot_get_project, depot_list_projects and the build/CI siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The idempotency note ('if an active project with the same name exists, that project is returned') implicitly tells the agent it can call this without a prior lookup, but there is no explicit when-to-use/when-not guidance or named alternative such as depot_get_project for read-only lookup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_get_buildGet a buildARead-onlyInspect
Get one container build's status, timestamps, duration, time saved and cached/total step counts. Depot: core BuildService/GetBuild.
| Name | Required | Description | Default |
|---|---|---|---|
| build_id | Yes | Build id, from depot_list_builds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety profile is covered. The description adds the exact data fields returned, which is useful transparency for an inspection call, but does not cover error behavior (e.g., unknown build_id) or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences: the first front-loads what is returned, the second locates the tool within Depot's service hierarchy. No wasted words, though the service-path sentence is somewhat redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param read-only tool with 100% schema coverage and annotations, the description is nearly complete. The only gap is the absence of any routing guidance to sibling build-inspection tools, which is minor given the clear field listing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is fully documented in the schema (including its source). The description adds no parameter-level detail beyond what the schema already provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (one container build) and enumerates the returned fields (status, timestamps, duration, time saved, step counts). This distinguishes it from depot_list_builds, but the description does not explicitly name that sibling or other alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description names the build_id and its source ('from depot_list_builds'), implying the typical workflow of listing then getting. But it does not state when to use this vs. depot_list_builds or the step-specific siblings (depot_get_build_steps, depot_get_build_step_logs).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_get_build_step_logsGet a build step's logsARead-onlyInspect
Get the log lines (message + timestamp) of one container-build step, identified by its digest from depot_get_build_steps. Depot: build.v1 BuildService/GetBuildStepLogs.
| Name | Required | Description | Default |
|---|---|---|---|
| build_id | Yes | Build id. | |
| page_size | No | Page size, 1-1000. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. | |
| project_id | Yes | Project id the build belongs to. | |
| step_digest | Yes | The step's digest, from depot_get_build_steps. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already declares the safety profile, so the bar is lower. The description usefully discloses that returned log lines carry message and timestamp, but says nothing about pagination behavior despite page_size/page_token parameters. Adds modest value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and immediately followed by the key dependency. No wasted words; the service identifier is tacked on at the end unobtrusively.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with full schema coverage and no output schema, the description adequately covers purpose, the digest dependency, and the shape of returned log entries. Pagination semantics remain a minor gap but the tool is well-specified overall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces that step_digest originates from depot_get_build_steps, but adds no syntax, format, or default details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: retrieving 'log lines (message + timestamp) of one container-build step'. It further distinguishes itself from sibling log tools by scoping to a container-build step and pointing at the digest source, depot_get_build_steps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the prerequisite workflow step explicitly — the digest comes from depot_get_build_steps — which tells an agent how to obtain the required identifier. It stops short of naming alternatives (e.g., depot_ci_get_job_logs) or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_get_build_stepsList a build's stepsARead-onlyInspect
List the BuildKit steps of a container build: name, digest, start/end time, cache state (CACHED / UNCACHED), any error, and whether it has logs. Find the failing step here, then pass its digest to depot_get_build_step_logs. Depot: build.v1 BuildService/GetBuildSteps.
| Name | Required | Description | Default |
|---|---|---|---|
| build_id | Yes | Build id, from depot_list_builds. | |
| page_size | No | Page size, 1-1000. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. | |
| project_id | Yes | Project id the build belongs to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description carries the rest: it discloses the returned fields, the CACHED/UNCACHED cache-state vocabulary, and the error/logs indicators. It also names the backing RPC (build.v1 BuildService/GetBuildSteps), though it doesn't mention pagination behavior on the read path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the field list front-loaded and the chaining instruction second; the RPC provenance tag at the end is low-value but brief. No wasted prose overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description's enumeration of returned fields is genuinely necessary and present. The only gap is that the read-path pagination flow (page_token chaining) isn't reinforced in prose, though the schema covers it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (project_id, build_id, page_size, page_token) are already documented in the schema. The description adds no parameter-level detail beyond referencing the step digest as an output, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the BuildKit steps of a container build') and enumerates the exact fields returned (name, digest, start/end time, cache state, error, logs flag), which cleanly separates it from depot_get_build and depot_get_build_step_logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete workflow: 'Find the failing step here, then pass its digest to depot_get_build_step_logs,' which routes the agent to the correct sibling. It lacks explicit when-not guidance or mention of the pagination flow, but the primary usage condition is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_get_projectGet a projectARead-onlyInspect
Get one Depot project — name, region, hardware size, cache policy, build timeout and autoscaling. Depot: ProjectService/GetProject.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Project id, from depot_list_projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safety profile, so the description is free to add value elsewhere — and it does by disclosing the concrete fields returned, which matters because no output schema exists. Minor boilerplate ('Depot: ProjectService/GetProject') adds little beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences that front-load the action and the payload before the trailing service identifier. Nothing is wasted except the 'Depot: ProjectService/GetProject' provenance string, which is marginal for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param read, annotations cover safety and the description compensates for the absent output schema by listing returned fields. It is essentially complete, with only error behavior (unknown id) left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter at 100% schema description coverage, and the schema itself explains it is a project id sourced from depot_list_projects. The description adds nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get one Depot project') and enumerates the fields returned (name, region, hardware size, cache policy, build timeout, autoscaling), so the agent knows exactly what it retrieves. The word 'one' implicitly separates it from the sibling depot_list_projects, though that sibling is never named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: fetching a single project by id is self-evident, and the schema's 'from depot_list_projects' hints at the discovery flow. There is no statement of when to prefer this over depot_list_projects or what happens for an unknown id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_get_usageGet organization usageARead-onlyInspect
Get the organization's Depot usage over a window: container builds per project (build count, minutes saved, minutes billed), GitHub Actions jobs per repo/workflow/runner (minutes elapsed and billed), storage, and agent sandboxes. Depot: UsageService/GetUsage.
| Name | Required | Description | Default |
|---|---|---|---|
| end_at | Yes | Window end, RFC 3339 (e.g. 2026-09-30T00:00:00Z). | |
| start_at | Yes | Window start, RFC 3339 (e.g. 2026-09-01T00:00:00Z). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description usefully adds that this is a cross-domain aggregate (builds, GHA jobs, storage, sandboxes) and gives the backing RPC path (UsageService/GetUsage), but says nothing about auth/permission scope, data latency, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence that leads with the verb and resource before the enumeration of outputs. The category list is long but each item carries information about what the response contains; the trailing 'Depot: UsageService/GetUsage' is minor filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does useful duty by naming the result categories the agent will receive. Read-only status comes from annotations and parameters are fully documented in the schema, so the only real gap is the absence of alternative-routing guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both date-time parameters are documented in the schema with RFC 3339 format and examples. The description only echoes this via 'over a window', so it adds no syntax or boundary-rule meaning beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the organization's Depot usage') plus the scoping dimension ('over a window'), then enumerates exactly what is aggregated: container builds per project, GHA jobs per repo/workflow/runner, storage, and agent sandboxes. That enumeration implicitly separates it from the narrower sibling depot_gha_get_analytics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer it is the aggregate usage-reporting call for a time window, and that start_at/end_at must bracket the window. No explicit when-to-use, when-not-to-use, or named alternative (e.g. depot_gha_get_analytics for a GHA-only view) is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_gha_get_analyticsGet GitHub Actions analyticsARead-onlyInspect
Get time-bucketed GitHub Actions analytics for Depot runners: job counts (successful/failed/cancelled), failure rate, elapsed and billable seconds, average duration — as totals plus hourly or daily buckets. Defaults to the last 30 days (max 90). Depot: GithubActionsService/GetGithubActionsAnalytics.
| Name | Required | Description | Default |
|---|---|---|---|
| jobs | No | Job display names to include. | |
| end_at | No | Window end, RFC 3339 (e.g. 2026-09-18T00:00:00Z). | |
| start_at | No | Window start, RFC 3339 (e.g. 2026-09-17T00:00:00Z). | |
| workflows | No | Workflow names to include. | |
| repositories | No | Connected repositories in owner/name format, e.g. ["acme/widgets"]. Empty = all. | |
| runner_labels | No | Runner labels to include, e.g. depot-ubuntu-24.04-8. | |
| workflow_paths | No | Workflow file paths to include, e.g. .github/workflows/ci.yml. | |
| aggregation_unit | No | Bucket size. Default: day. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description still adds real behavioral context beyond the annotations: the default 30-day window, the 90-day maximum, and the fact that results come as totals plus hourly/daily buckets. It omits pagination or result-size limits, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the verb and the metric list, with the window default and cap following. The trailing internal identifier 'Depot: GithubActionsService/GetGithubActionsAnalytics' adds little for an agent selecting a tool, minor clutter that keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, zero-required read tool with no output schema, the description does meaningful work: it enumerates the returned metrics and the bucket granularity, and states the default and maximum windows. It is nearly complete, missing only guidance on result limits or filtering semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all eight parameters are already documented in the schema. The description only echoes the bucketing option ('hourly or daily buckets') and the date window default, adding marginal meaning beyond the structured fields; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Get time-bucketed GitHub Actions analytics') and enumerates the exact metrics returned (job counts, failure rate, elapsed/billable seconds, average duration), so the agent knows what it produces. It does not explicitly differentiate itself from analytics-adjacent siblings like depot_get_usage, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The default window (last 30 days, max 90) tells the agent how the tool behaves if it omits parameters, which implies a typical usage context. However, there is no explicit when-to-use statement and no routing to alternatives such as depot_gha_list_jobs or depot_gha_search_logs, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_gha_list_jobsList GitHub Actions jobs on Depot runnersARead-onlyInspect
List GitHub Actions jobs that ran on Depot-managed runners, newest first: status, conclusion, runner label, requested CPU/memory, duration, billable time and GitHub URL. Defaults to the last 30 days (max window 90 days). Depot: GithubActionsService/ListGithubActionsJobs.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Case-insensitive text matched against job and workflow names; numeric values also match exact GitHub job/run ids. | |
| end_at | No | Window end, RFC 3339 (e.g. 2026-09-18T00:00:00Z). | |
| start_at | No | Window start, RFC 3339 (e.g. 2026-09-17T00:00:00Z). | |
| statuses | No | Job states to include. | |
| page_size | No | Page size, 1-250. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. | |
| conclusions | No | Job conclusions to include, e.g. ["failure"]. | |
| repositories | No | Connected repositories in owner/name format, e.g. ["acme/widgets"]. Empty = all. | |
| runner_labels | No | Runner labels to include. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so safety is already covered, and the description adds real behavioral constraints: an implicit 30-day default window and a hard 90-day maximum, plus newest-first ordering. It does not discuss pagination limits, result caps, or authentication requirements, so it stops short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler, front-loading verb, resource, and scope before the returned fields and the date-window constraints. The trailing service identifier is a compact provenance note rather than padding; nothing here could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With nine optional parameters and no output schema, the description usefully substitutes for the missing return schema by enumerating the reported fields, and it covers the window defaults that the schema omits. It leaves pagination behavior and empty-result semantics unstated, though the page_token/page_size params partially cover paging.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so 3 is the baseline, but the description adds information the schema lacks: the default date range and the 90-day maximum window, which governs how start_at/end_at must be set. It does not add meaning for query, statuses, conclusions, repositories, runner_labels, or the paging parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a precise verb and resource ('List GitHub Actions jobs') and adds a hard scope qualifier ('that ran on Depot-managed runners'), which separates it from run-level siblings like depot_ci_list_runs. It even enumerates the returned fields (status, conclusion, runner label, CPU/memory, duration, billable time, URL), so an agent knows exactly what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies operating context ('defaults to the last 30 days', 'max window 90 days', 'newest first'), which tells the agent how to call it, but never states when to prefer it over depot_ci_list_runs, depot_gha_get_analytics, or depot_gha_search_logs. Usage is implied by the resource scope rather than explicitly routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_gha_list_recommendationsList runner-size recommendationsARead-onlyInspect
List Depot's runner-size recommendations from job CPU and memory metrics: SIZE_UP when average CPU or memory reaches 90%, SIZE_DOWN when both averages stay ≤30% and peaks ≤70%, with current vs suggested size and per-minute price. Defaults to the last 30 days. Depot: GithubActionsService/ListGithubActionsRecommendations.
| Name | Required | Description | Default |
|---|---|---|---|
| jobs | No | Job display names to include. | |
| end_at | No | Window end, RFC 3339 (e.g. 2026-09-18T00:00:00Z). | |
| start_at | No | Window start, RFC 3339 (e.g. 2026-09-17T00:00:00Z). | |
| page_size | No | Page size, 1-250. | |
| workflows | No | Workflow names to include. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. | |
| repositories | No | Connected repositories in owner/name format, e.g. ["acme/widgets"]. Empty = all. | |
| runner_labels | No | Runner labels to include, e.g. depot-ubuntu-24.04-8. | |
| workflow_paths | No | Workflow file paths to include, e.g. .github/workflows/ci.yml. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds real behavioral context beyond that: the exact thresholds that produce SIZE_UP (avg CPU or memory ≥90%) versus SIZE_DOWN (both averages ≤30%, peaks ≤70%), plus the default 30-day lookback and the fields returned. Pagination/rate-limit behavior is left unmentioned, which keeps it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose, thresholds, and default window are packed into one front-loaded sentence with no filler, so scanning yields the essentials immediately. The trailing internal API path ("Depot: GithubActionsService/ListGithubActionsRecommendations.") is decorative and doesn't help tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partly compensates by describing the returned fields (current vs suggested size, per-minute price). All 9 parameters are documented in the schema, and readOnlyHint covers safety, so the omission of pagination details and per-filter semantics is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning the schema does not: the 30-day default window gives an agent a concrete interpretation of the otherwise optional start_at/end_at parameters, removing the need to guess the temporal scope when filters are omitted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List Depot's runner-size recommendations") and names the governing data source (job CPU and memory metrics), so an agent can distinguish it from adjacent listing tools such as depot_gha_list_jobs or depot_gha_get_analytics. The recommendation vocabulary (SIZE_UP/SIZE_DOWN) further pins down the resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the description explains what triggers each recommendation type and that the window defaults to 30 days, which tells an agent when the output is relevant. But it never states explicit when-to-use or when-not-to-use conditions, nor names an alternative tool for raw metrics versus recommendations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_gha_search_logsSearch GitHub Actions logsARead-onlyInspect
Search GitHub Actions log lines from Depot runners, newest match first (job, workflow, step, timestamp, body, lineId). Defaults to the last hour (max window 30 days). Filter by repository, workflow, runner label, action or compute id. Depot: GithubActionsService/SearchGithubActionsLogs.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Text to match in log bodies, e.g. "timeout". Empty returns unfiltered lines. | |
| end_at | No | Window end, RFC 3339 (e.g. 2026-09-18T00:00:00Z). | |
| actions | No | Actions to include, e.g. actions/checkout. | |
| start_at | No | Window start, RFC 3339 (e.g. 2026-09-17T00:00:00Z). | |
| page_size | No | Page size, 1-1000. | |
| workflows | No | Workflow names to include. | |
| compute_id | No | A single Depot compute id. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. | |
| repositories | No | Connected repositories in owner/name format, e.g. ["acme/widgets"]. Empty = all. | |
| runner_labels | No | Runner labels to include. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already covering the safety profile, the description adds genuinely new behavior: results are newest-match-first, the default window is one hour, and the maximum window is 30 days. It omits any note on result limits or how the cursor behaves beyond the schema's page_token wording, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler; the search scope and ordering lead, followed by the time-window default/limit and filter dimensions. Every clause carries information an agent would otherwise have to infer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by naming the returned fields (job, workflow, step, timestamp, body, lineId), plus ordering, default window, and max window. With readOnlyHint covering safety and the schema covering all 10 parameters, nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameter meaning is already fully documented. The description restates the filterable dimensions (repository, workflow, runner label, action, compute id) and the time-window bounds, which reinforces but does not add new syntax or semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search GitHub Actions log lines from Depot runners') plus the ordering and the returned fields, so an agent immediately knows this is a text-search over GHA log lines rather than a fetch of a whole job/build log. Combined with the sibling set (depot_ci_get_job_logs, depot_get_build_step_logs), the scope is cleanly separable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: defaults to the last hour, max 30-day window, and the dimensions you can filter by (repository, workflow, runner label, action, compute id). It does not name an alternative tool or say when not to use this one, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_list_buildsList builds for a projectARead-onlyInspect
List container builds in a project, newest first: status (STATUS_RUNNING / FAILED / SUCCESS / ERROR / CANCELED), timestamps, build duration, time saved by cache, and cached vs total steps. Depot: core BuildService/ListBuilds.
| Name | Required | Description | Default |
|---|---|---|---|
| page_size | No | Page size, 1-100. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. | |
| project_id | Yes | Project id, from depot_list_projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, so safety is covered. The description adds real value beyond that by disclosing the returned fields (status values, timestamps, duration, cache time saved, cached vs total steps), which matters because there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence with the resource and ordering front-loaded, followed by a short backend-reference trailer. Efficient and well-ordered, though the 'Depot: core BuildService/ListBuilds' clause is largely decorative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned fields and status values, and pagination semantics live in the schema. Complete enough for an agent to list builds correctly; only the when-to-use routing is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so project_id, page_size and page_token are already documented in the schema, including that project_id comes from depot_list_projects. The description adds no parameter-level detail beyond reinforcing project scoping, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list container builds in a project) plus an ordering guarantee (newest first), which cleanly separates it from the singular depot_get_build. It stops short of naming a sibling tool to route against, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope phrase 'in a project' implies you need a project and that this is the collection-level listing, but there is no explicit when-to-use guidance, no mention of depot_get_build/depot_get_build_steps for single-build or step-level detail, and no exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_list_projectsList projectsARead-onlyInspect
List the organization's Depot projects. A project is an isolated container-build cache (region, hardware, cache policy). Use the returned projectId with the build and registry tools. Depot: ProjectService/ListProjects.
| Name | Required | Description | Default |
|---|---|---|---|
| page_size | No | Page size, 1-250. | |
| region_id | No | Only projects in this region or cloud connection, e.g. us-east-1. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes this as a safe read, so the bar is lower. The description adds useful framing (what a project contains and that its projectId feeds other tools), but says nothing about pagination behavior or result ordering beyond what the schema's page_token already implies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, then the domain definition, then the practical follow-up. The trailing 'Depot: ProjectService/ListProjects' internal reference earns its place only marginally but costs little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, zero-required-parameter list tool with full schema coverage and no output schema, the description is nearly sufficient: purpose, entity definition, and how to use the result are all present. Only pagination-flow behavior is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so page_size, region_id, and page_token are fully documented in the schema. The description adds no syntactic or semantic detail about any of them, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the organization's Depot projects') and then defines what a project is ('an isolated container-build cache with region, hardware, cache policy'). This distinguishes it from the singular sibling depot_get_project and from depot_list_builds / depot_list_registry_images without any schema lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear downstream context: 'Use the returned projectId with the build and registry tools.' That tells the agent why and when to call it, but it never names alternatives (e.g., depot_get_project for a single project) or states any exclusion condition, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
depot_list_registry_imagesList registry imagesBRead-onlyInspect
List the images stored in a project's Depot Registry: tag, digest, pushedAt and sizeBytes. Depot: build.v1 RegistryService/ListImages.
| Name | Required | Description | Default |
|---|---|---|---|
| page_size | No | Page size, 1-100. | |
| page_token | No | Cursor from the previous response's nextPageToken. Omit for the first page. | |
| project_id | Yes | Project id, from depot_list_projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe, non-mutating read. The description adds value by disclosing the payload fields (tag, digest, pushedAt, sizeBytes) in the absence of an output schema, but says nothing about paging behavior beyond what the params imply or about ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences; the scope and returned fields are front-loaded with no filler. The trailing 'Depot: build.v1 RegistryService/ListImages' is boilerplate but low-cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with a fully documented schema and readOnlyHint covering safety, the description supplies the one thing structured data lacks: the shape of the returned records. Little else is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three parameters are documented in the schema (page_size range, page_token cursor semantics, project_id provenance). The description adds no parameter meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('images stored in a project's Depot Registry') and enumerates the returned attributes (tag, digest, pushedAt, sizeBytes). The registry scope is distinct from CI/build/project siblings, though it never names a sibling to differentiate against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus depot_list_builds or the CI listing tools. The only implied usage is 'list images in a project,' and there are no prerequisites, exclusions, or alternative-routing hints.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
22 tool updates
- First observed
depot_ci_cancel_run - First observed
depot_ci_get_failure_diagnosis - First observed
depot_ci_get_job - First observed
depot_ci_get_job_logs - First observed
depot_ci_get_run_status - First observed
depot_ci_list_runs - First observed
depot_ci_list_workflows - First observed
depot_ci_retry_failed_jobs - First observed
depot_ci_retry_job - First observed
depot_create_project - First observed
depot_get_build - First observed
depot_get_build_step_logs - First observed
depot_get_build_steps - First observed
depot_get_project - First observed
depot_get_usage - First observed
depot_gha_get_analytics - First observed
depot_gha_list_jobs - First observed
depot_gha_list_recommendations - First observed
depot_gha_search_logs - First observed
depot_list_builds - First observed
depot_list_projects - First observed
depot_list_registry_images
Related MCP Connectors
Trigger and inspect Codemagic CI builds, apps and artifacts.
Check if a repo will deploy, plan a deploy to your own cloud, and see status, logs and redeploys.
Debug background jobs — runs, traces, spans, schedules, queues and deployments.
Triage failing GitHub Actions jobs and see what self-heal repaired, in natural language.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides coding agents read-only access to Depot's CI failure diagnoses, container build forensics, run history, cache effectiveness, registry contents, and usage data through MCP tools.2848 npm1-
- AlicenseAqualityAmaintenanceGoverned CI/CD operations for self-hosted GitLab and Gitea — pipeline-failure, runner, artifact-bloat, and stale-branch RCA, with unbypassable audit logging (MCP + CLI), budget/runaway guards, dry-run, and undo/rollback.28MIT
- AlicenseAqualityCmaintenanceProvides tools to analyze and debug GitHub Actions CI failures, including summarizing failures, detecting flaky tests, and suggesting fixes.108 npm1ISC
- FlicenseNot gradedqualityDmaintenanceEnables interaction with Buildkite CI/CD to list organizations, pipelines, builds, jobs, and logs, as well as retry jobs.4-
Glama MCP Gateway
Add one secure layer between your agents and this server.