Skip to main content
Glama
gagarin-cloud

gagarin MCP server

Official

mcp

mcp.gagarin.cloud — gagarin's API, as tools an agent can call.

npm install
npm run build
npm test                 # builds, then node --test over dist/
npm start                # the HTTP server on :8080
npm run stdio            # the same tools over a pipe

The Model Context Protocol is how an agent finds and calls a tool it was not built with. This is gagarin's, and it exists so that adding gagarin to a coding agent is a URL rather than an install — the one artefact that makes the platform installable rather than merely documented.

Two ways in, one implementation:

remote

https://mcp.gagarin.cloud/mcp, streamable HTTP, credential in the Authorization header — put there by OAuth sign-in or by hand

local

npm run stdio from a clone of this repository, credential from GAGARIN_TOKEN or the file gg login wrote — for developing on this server, not a way to install it

Signing in

Remote, with OAuth. Give the client the URL and nothing else:

https://mcp.gagarin.cloud/mcp

Claude, ChatGPT and Claude Code prompt for sign-in when they connect: the client opens a browser and the human signs in with GitHub or Google. A request with no credential answers 401 with a WWW-Authenticate header pointing at /.well-known/oauth-protected-resource/mcp, which names api.gagarin.cloud as the authorization server. This server decides nothing about a token itself; the engine does, on every call. GAGARIN_MCP_ORIGIN and GAGARIN_ISSUER override the two public names for a development setup; they are separate from GAGARIN_API, which in the cluster is the in-cluster Service and no address a client could sign in at.

Remote, with a credential in a header. For a client that cannot sign in over OAuth, or a machine that should not: a credential from gg login or gg creds create.

{
  "mcpServers": {
    "gagarin": {
      "type": "http",
      "url": "https://mcp.gagarin.cloud/mcp",
      "headers": { "Authorization": "Bearer <your gagarin credential>" }
    }
  }
}

Local, over stdio. npm run stdio from a clone runs the same tools over a pipe, reading the file gg login wrote, or GAGARIN_TOKEN if it is set. It is here to develop against, and is deliberately not published to npm: the point of this server is that adding gagarin to an agent is a URL and not an install, and a package on somebody's laptop is a second copy of src/tools.ts that goes stale the day a tool changes.

A credential that has expired or been revoked is caught at the door too, because a client signs in again only on an HTTP 401: every POST asks the engine /v1/whoami with the caller's token first, and the engine's 401 becomes a 401 with error="invalid_token". One extra in-cluster call per request, and no cache — a remembered token is a stored token, and this server stores none. Any other failure there (the engine unreachable, a 5xx) is not a sign-in problem, so the request goes on and each tool reports it with the engine's own code.

Related MCP server: Scout MCP Server

What is here

path

what it is

src/api.ts

the whole of this server's contact with gagarin: one request, one error envelope

src/tools.ts

every tool — each one a path, a shape, and the rule a caller needs before using it

src/server.ts

what a client is told on connect, and the gagarin://guide resource — including how memory is used and the .gagarin.json note that says which project a repository is

src/app.ts

mcp.gagarin.cloud: stateless streamable HTTP, one server per request, and the OAuth protected-resource metadata

src/http.ts

the listener, its configuration and its drain

src/stdio.ts

the same tools over a pipe, for developing against from a clone

src/credentials.ts

reads the credential file gg login wrote; never writes one

Dockerfile

the image mcp.gagarin.cloud runs. Its build stage runs the tests

Project memory

Every project carries a memory: small, durable facts an agent saves about a codebase — a decision and why, a convention, a gotcha — so the next session does not rediscover them. It is a built-in of the project, like its registry, and these tools are the only way to reach it: the memory service has no address of its own, no gg command and no console page. The engine decides who may read (viewer) and who may write (editor), on every call, like everything else.

tool

what it does

memory_briefing

what is known about a project, packed to a token budget — the first call on a project

memory_search

hybrid semantic + keyword search; no query browses by rank

memory_get

memories in full, by id

memory_related

walk the links out from one memory

remember

save one fact; a near-duplicate is refused with the memories it collided with

memory_update

change fields, pin, or archive

memory_link / memory_unlink

associate two memories, or stop

These answer the memory service's own text — a plain-text rendering packed to the budget — rather than the JSON around it. That is a narrower case of the rule below, not an exception to it: the rendering is still the engine's, and returning both would spend twice the tokens the service exists to save. A body without text comes back as JSON like any other answer.

It is a translator, and nothing else

It holds no credential of its own, has no database, and makes no decision the API does not make. Every tool is one call to api.gagarin.cloud carrying the caller's own bearer token, and every refusal is the engine's refusal passed through unedited.

That is the property everything else rests on: possessing this server grants nothing at all. It is what makes it safe to put a public endpoint in front of a single write gate, and nothing here may be changed in a way that weakens it.

Two consequences worth stating, because both look like omissions:

  • Answers are the engine's JSON, not a summary of it. Every service, ledger line and connection already carries a sentence written by the engine, precisely so a terminal and a dashboard cannot describe the same row differently. A third renderer here would be a third opinion to keep in step.

  • Errors are [code] message with a hint: line, which is what gg prints. One format, so an agent that has read the gagarin skill recognises what comes back here without being taught a second one. When the engine sends more than the envelope — a memory_duplicate carries the near-duplicates — that is appended after the hint rather than dropped.

What it deliberately cannot do

Four things need a machine, and offering them here would produce failures that read like platform faults:

  • Build and push an image. gg ship — build, push and deploy fused — shells out to docker where the source is. deploy here runs an image that is already in gagarin's registry: a tag CI pushed, or a restatement. Getting one there is the CLI's job.

  • Open a tunnel. gg connect binds a local port.

  • Wait for a job to finish. run submits and returns a revision, like every other write here. gg run blocks until the run ends and exits with the script's own exit code, which is what a pipeline wants; over MCP you poll status for the phase and the code.

  • Click an approval. That is a human with an inbox, by design.

There is also no list of resource types, sizes or scopes in this repository. The engine owns those and its refusals name them; a z.enum here would be a second list in a second repository, wrong on the day a type is added — a mistake this codebase has already made once, in gg, about this exact family of values.

Where it runs

In the Kubernetes cluster, next to the control plane, and not on Vercel beside the site and the console. Those two live outside Scaleway because their job is to still be there and say the platform is down. This one has nothing whatever to say when the API is unreachable; it is that API in another shape, so it belongs next to it and reaches it over the in-cluster Service rather than hairpinning out through the load balancer.

The pod runs as a ServiceAccount with no RBAC and no mounted token, on a read-only root filesystem. It needs none of them.

Every push to main builds the image and rolls it out — .github/workflows/deploy.yml. The credentials that can do that are secrets of the production environment, which only main may use, so a pull request never runs next to them. Merging to main is deploying.

Adding a tool

One server.registerTool call in src/tools.ts: a name, a description saying the rule a caller most needs, a zod shape, and one api.call. Then a test in src/tools.test.ts asserting the path, the method and the body — those are the only things it is possible to be quietly wrong about, and a deploy that PUTs to the wrong path answers 404 and reads like a missing service.

Do not add a tool for something the API does not do. This file has no business being the place a feature appears first.

Available Tools

35 tools
add_domainPut a service on the internetA
Idempotent
Inspect

With no domain, hands out gagarin's own generated address and the certificate is already held. With one, claims that name — and the answer says what DNS record the owner has to add before it can be issued. Both are idempotent; restating one repairs its ingress. A service is private until this call and a deploy can neither give an address nor take one away.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNoa hostname you control, e.g. shop.example.com. Absent asks for the generated one.
projectYesproject name or id
serviceYesservice name, unique within the project

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description reveals idempotent ingress repair, certificate and DNS issuance prerequisites, and the privacy/state change of the service. These are material behaviors not otherwise visible in annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each adding a distinct fact: the no-domain case, the custom-domain case, idempotent repair, and public/private state behavior. The key distinction is front-loaded with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-mode tool with no output schema, the description covers invocation modes, idempotency, DNS follow-up, and state implications. Nothing essential is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description enriches the optional domain parameter: absence means a generated address with an already-held certificate, while presence triggers a custom name claim and DNS requirement. It adds meaning beyond the schema's straightforward hostname description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise action: put a service on the internet either via a generated address or a claimed custom domain. It also conveys the before/after state, which distinguishes it from siblings like remove_domain and deploy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear when-to-use context: omit the domain for a generated address, provide a controlled hostname for a custom one, and note that a deploy cannot add or remove an address. It does not explicitly name excluded alternatives, but the conditional guidance covers the main decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_resourceProvision a managed resourceA
Idempotent
Inspect

Something gagarin runs on your behalf — a database, a cache, a vector store — or an external, a row that runs nothing and only publishes values. You name it and say how big; the platform decides the rest. If a type exists, use it rather than deploying your own; the refusal from an unknown type names the ones there are. Nothing reaches it until a service declares it in deps. Use an external for third-party credentials AND for shared configuration — a feature flag, a log level, a region, an API base URL — anything more than one service reads, or that should change without a deploy. An env passed to a deploy is a copy, so two services sharing a setting is two copies that can disagree and two deploys to change it; an external is one row, changed with rotate_resource (set/unset for one key), undone with rollback, and every holder is restarted for it. Config owned by a single service stays in its deploy env, because that is the half a service rollback restores. The decisive difference when you do not have the user's env file: a service's environment can only be changed by deploy, which replaces it wholesale, so changing one variable means having all of them. An external takes a single key. If a user asks you to change a setting and you were not given their environment, an external is the answer — and if that setting currently lives in a deploy env, say so and offer to move it. Keys are published under the resource name — config holding LOG_LEVEL publishes CONFIG_LOG_LEVEL — so name the resource for how the application wants to read it.

ParametersJSON Schema
NameRequiredDescriptionDefault
envNothe values an `external` publishes to whatever declares it. Refused for every other type, whose credentials are gagarin's to mint rather than yours to choose.
sizeNothe same envelope word a service takes. Changeable later.
typeYeswhat to provision, e.g. postgres, valkey, qdrant or external
projectYesproject name or id
resourceYesresource name, unique within the project
storage_gbNohow big its volume may get. Fixed at creation, like every volume.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though annotations already mark this as non-read-only and idempotent, the description adds substantial behavioral context: nothing reaches the resource until a service declares it in `deps`, externals publish keys under the resource name, unknown types produce refusals that name valid types, and changing an external restarts every holder. This goes well beyond what the annotations convey and gives the agent a realistic model of the tool's effects. No contradiction with annotations is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but nearly every sentence earns its place by explaining a decision or behavior that is not obvious from the schema or annotations. It uses bolded directives to front-load the most actionable guidance, though the opening conceptual framing could be tighter. It is dense rather than bloated, and the structure supports scanning for key rules.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six parameters, two resource categories, and several interacting operational concerns, the description is remarkably complete. It covers what the tool does, when to use it over siblings, how to name resources, how externals differ from deploy env, how changes propagate, and what to do when the user's environment is unavailable. No output schema exists, but the tool's return value is not critical to correct invocation, so the absence does not create a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters. The description still adds value by explaining the conceptual role of `env` (what an external publishes, refused for other types), the naming convention that ties `resource` to published keys, and the meaning of `size` as an envelope word. This exceeds the baseline for fully covered schemas, though it does not deeply elaborate every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action — provisioning a managed resource — and immediately enumerates the kinds of things it covers (database, cache, vector store, external). It also distinguishes the `external` subtype from managed resources, so the tool is clearly separated from anything that simply creates a database. The purpose is unmistakable even before reading the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use and when-not-to-use guidance, especially for `external` versus deploy env: shared settings and third-party credentials should be externals, single-service config should stay in deploy env, and if the user's env file is absent, an external is the answer. It also names alternatives like `rotate_resource`, `rollback`, and `deploy`, directly routing the agent away from this tool when a different operation is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backup_resourceTake a backup nowBInspect

A snapshot at this moment, on top of the nightly ones. Rate limited, so do not call it in a loop.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
resourceYesresource name, unique within the project

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the description's role is additive. It contributes the rate-limit warning and the 'on top of nightly ones' scoping, which are useful. However, it does not disclose whether the call is synchronous or asynchronous, whether it generates a backup ID or how to track completion, or any cost/retention implications. With annotations present, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. It front-loads the purpose ('A snapshot at this moment'), then adds the operational constraint ('Rate limited, so do not call it in a loop'). Every word earns its place, and it is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter action, the description covers the what and a key operational trait, but it leaves gaps: there is no mention of how the agent will know the backup succeeded (no output schema), no indication of whether the operation is synchronous, and no hint that progress can be checked via the sibling 'backups' tool. It is minimally sufficient but not complete enough to fully guide an agent without prior knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides descriptions for both parameters ('project name or id' and 'resource name, unique within the project') at 100% coverage, so the description adds no parameter-specific meaning. It does not reinforce that project/resource identify the target of the backup, but the schema already carries that burden. The baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly conveys that this tool creates an immediate snapshot in addition to nightly backups. The phrase 'A snapshot at this moment' expresses the core action, and 'on top of the nightly ones' distinguishes it from scheduled backup behavior. However, it is phrased as a noun rather than an explicit verb+resource, so it slightly misses top marks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is 'Rate limited, so do not call it in a loop,' which is a usage constraint but not a when-to-use directive. It does not mention alternatives like the sibling 'backups' (for listing) or 'restore_resource' (for restoring), nor does it state when an ad-hoc snapshot is appropriate. The agent is left to infer the decision context from the tool name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backupsWhat has been backed upA
Read-only
Inspect

Every stored backup of a resource, newest last. Keys are UTC timestamps, so they sort chronologically, and one is what restore_resource takes for an exact point.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
resourceYesresource name, unique within the project

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint=true and openWorldHint=true, indicating a safe read operation. The description adds useful behavioral context: the output is ordered by timestamp (newest last) and that timestamps are keys. It doesn't contradict annotations. The bar is lower because annotations exist, and the description adds value by describing the output ordering and the relationship to restore. However, it doesn't detail the output format beyond that, so a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, both of which convey essential information. The first sentence is efficient and informative, and the second adds crucial context about the timestamps and their relation to restore_resource. No waste, and the key behavioral detail (newest last) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with high schema coverage, the description is quite complete. It tells the agent what the output is, the ordering, and how it relates to restore. It lacks details about the output structure (e.g., fields of each backup entry), but since there's no output schema and the context is simple, this is a minor gap. An agent can confidently call this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (project and resource) are well-documented in the schema. The description adds minimal parameter semantics: it implies that 'resource' is the resource in question and that the output is scoped to that resource. It doesn't add syntax or format details, which is acceptable given high schema coverage. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states what the tool does: lists stored backups of a resource. The phrase 'Every stored backup of a resource' is specific and distinguishes it from backup creation and restoration. However, it could be more explicit that it is a list/read operation, though the name 'backups' and the context imply it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides good usage context by noting that the keys are UTC timestamps and that one is what restore_resource takes for an exact point. This hints at when to use this tool (before restoring) and aids selection. However, it doesn't explicitly state when NOT to use it or name alternatives like restore_resource or backup_resource directly, but the reference to restore_resource provides enough guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

billingWhere the account standsA
Read-only
Inspect

Balance, burn rate and runway — how long what is running now can keep running. A suspended account refuses every deploy, and a runway measured in hours is worth telling your human about before it becomes an outage.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, and the description adds meaningful behavioral context: the tool reports current financial viability, and it flags operational consequences of suspension or near-term runway exhaustion. It does not specify the exact response shape, but for a read-only status tool the core behavior is sufficiently transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the primary outputs (balance, burn rate, runway) and the second sentence earns its place by explaining why the result matters for agent prioritization.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only status tool, the description covers what the agent can expect conceptually and what actions may follow from the result. It could be more explicit about exact return units or fields, but the low complexity and clear scope make it largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema description coverage is 100%, so the schema fully defines the input surface. Per the rubric, this earns a baseline of 4; no parameter-level explanation is needed or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (account billing standing) and the key data points returned: balance, burn rate, and runway. It avoids tautology, but it is phrased as a set of concepts rather than a crisp verb+resource, and it does not explicitly position itself against the sibling billing_history tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when the tool matters: a suspended account blocks deploys, and a short runway warrants urgent human escalation. However, it does not explicitly state when to prefer this tool over alternatives like billing_history or platform_health, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

billing_historyThe ledgerA
Read-only
Inspect

The rows the balance is folded from: metered usage every quarter hour, top-ups, credits and adjustments. Each carries a sentence written by the engine and an amount already rendered, so nothing here needs re-computing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint; the description adds that rows are pre-rendered with engine-written sentences and amounts, and that no recomputation is needed. This clarifies the data is a read-only, finalized history beyond what the annotations already communicate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences deliver the core purpose, the item types, and the key behavior (pre-rendered amounts). The information is front-loaded in the opening phrase and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only history view, the description is sufficient: it covers what rows exist, what fields they contain, and that no computation is required. The annotations handle the read-only and open-world aspects, and no output schema exists to document.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema is empty, so there are no parameters to explain; the baseline for zero parameters is 4. The description's content details are consistent with the absence of inputs, and no additional parameter guidance is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource as the ledger of billing rows and enumerates the categories (metered usage, top-ups, credits, adjustments). It is clear but lacks an explicit verb such as 'list' or 'retrieve,' and it does not name a sibling to differentiate from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It frames the tool as the source from which the balance is folded, so an agent can infer this is the place to inspect billing line items rather than the aggregate `billing` view. The 'nothing here needs re-computing' line also tells the agent the amounts are final and can be used directly. There is no explicit when-not-to-use statement, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_credentialMint a credential for CIAInspect

Issues a second credential from this one, for a pipeline that has no inbox to approve anything with. It can deploy and nothing else, it expires, and it cannot mint another — none of which are parameters. The secret comes back once and is never readable again, so put it straight into the secret store it is for. A credential that was itself minted cannot call this.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYeswhat you will read months from now when deciding whether this is still wanted, e.g. "github actions: acme/web"
expires_in_daysNo1 to 365. Absent takes the default.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations, disclosing that the new credential is deploy-only, expires, cannot mint another, returns its secret only once, and cannot be used to call this tool again if it was itself minted. These are critical behavioral traits not visible in the schema or annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences front-load the purpose and then compactly convey the key behavioral constraints)Skip. Every sentence contributes essential context without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a credential-minting tool with no output schema, the description covers the essential operational facts: what the credential can do, its expiry, the one-time secret retrieval, and the caller restriction. An agent has enough context to invoke it correctly and avoid misuse.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents both parameters fully, so the baseline is 3. The description adds that the credential's restrictions are not parameters, but this is slightly ambiguous given the existence of 'expires_in_days' as a parameter. It does not meaningfully enrich parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Issues a second credential from this one' for CI pipelines. It is clearly distinct from sibling tools like 'credentials' or 'revoke_credential' by describing a restricted, delegating credential rather than listing or revoking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit use case: 'for a pipeline that has no inbox to approve anything with.' It also adds restrictions like 'cannot call this' for minted credentials and describes what the new credential can do, but it does not name alternative tools or compare itself to them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_projectCreate a projectBInspect

A project is the unit of naming, access and billing: everything else lives inside one. Returns its id, which is what image paths are built from.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes2-30 chars, lowercase letters, digits and hyphens, starting with a letter

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal that this is a mutating, non-idempotent operation, so the description only needs to add context beyond that. It does add the useful detail that the tool returns the new project's id and that this id is the basis for image paths. However, it does not disclose side effects such as billing activation, uniqueness behavior, or failure cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, efficient, and has no filler; both sentences contribute meaning. It loses a point because the first sentence is definitional rather than action-oriented, and the second sentence is a fragment that relies on the title for its subject, but overall it is well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-required-parameter creation tool with no output schema, the description covers the key operational facts: what a project is, how the result is returned, and what that result is used for. It is not fully complete because it never explicitly states the create action or describes error/edge-case behavior, but those are less critical at this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage for the single parameter 'name' is 100%, including a detailed format description. The tool description adds nothing specifically about how the name parameter is validated or used, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title and tool name clearly indicate the create action, but the description itself never explicitly says 'creates a project.' It does add useful semantic context about what a project is and that the returned id is used to build image paths, which helps an agent understand the resource. It does not name or differentiate against sibling tools like projects or destroy_project, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance about when to use this tool versus alternatives, no prerequisites, and no mention of sibling tools such as projects or destroy_project. The phrase 'everything else lives inside one' implies that a project must exist first, but this is not framed as actionable usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

credentialsWhat has access to this accountA
Read-only
Inspect

Every credential, what it may do, when it was last used and when it expires. The one making this very call is marked current.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the tool as read-only, and the description adds meaningful detail about what the response contains, including capabilities, last-used timestamps, expiration, and the caller-specific `current` marker. It does not repeat annotation data and offers behavioral context that helps the agent interpret results, though it omits potential size or rate considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences carry substantial information with no filler. The first sentence enumerates the output fields, and the second provides a distinctive behavioral detail about the current credential, making every clause valuable and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description adequately conveys the scope and contents of the result: all credentials, permissions, usage recency, expiry, and the caller's marker. It could more explicitly state that the return is a list of credentials, but 'Every credential' strongly implies this, so essential information is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description adds no parameter-specific detail, but none is needed since the input schema is already complete and empty. The focus on return content is appropriate for a no-argument tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool enumerates every credential with its permissions, last-used time, expiry, and marks the caller's credential as `current`. The title reinforces that it answers 'what has access to this account,' making its purpose distinct from credential lifecycle tools like create_credential and revoke_credential.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes it obvious that this is the tool for auditing or inspecting all credentials, not just the caller's identity. It does not explicitly name when-not-to-use alternatives, but the 'current' marking implicitly differentiates it from whoami and credential management tools, providing clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deployRun an image that is already in the registryA
Idempotent
Inspect

Declares what a service should be. The image must already be in gagarin's own registry under this project — gagarin runs nothing else — and this server cannot put it there: building and pushing need docker and the source, so they happen on a machine, with gg ship (build, push and deploy in one) or gg build + gg push in CI. Use this tool to deploy an image that exists: a new tag CI pushed, or a restatement. For an image that runs to completion rather than listening — a migration, a backfill — that is run, and the two cannot be swapped: deploying over a job is refused not_a_service. Three rules the shape of this call depends on. Env is replaced wholesale, so restate every variable on every deploy — the domain, the size and the dependencies all survive a deploy that forgets to mention them. That also means you cannot change one variable without holding them all: if you were not given the environment, do NOT reconstruct it from history and redeploy — that drops anything you misread and pulls every secret the service holds through this conversation. Put the value in an external resource instead, where rotate_resource changes one key on its own. A volume must be restated too, and for the opposite reason: it cannot be changed, so a service that has one is refused volume_immutable unless volume_path and volume_size_gb come back exactly as status reports them. And a deploy neither gives an address nor takes one away: that is add_domain. A service is private until it has one, and a private service is reachable only by services that declared they need it.

ParametersJSON Schema
NameRequiredDescriptionDefault
envNothe complete environment. Absent empties it — this field is not merged.
depsNoservices and resources this one may reach, added to whatever it already declares. Adds and never removes — use `set_deps` to withdraw one. Here so a service that needs a database can be deployed holding its credentials from the first pod.
portYesthe TCP port the container listens on
sizeNoCPU and memory envelope. Absent keeps whatever it already had.
imageYesfull reference in gagarin's registry, e.g. registry.gagarin.cloud/<project-id>/web:v3. `whoami` gives the registry host and `projects` the id.
digestNosha256:... as docker push reported it, pinning the exact image
projectYesproject name or id
serviceYesservice name, unique within the project
volume_pathNoabsolute directory inside the container that survives a restart. Fixed at the deploy that creates the service and never changeable — which means every later deploy of that service must repeat it unchanged, or be refused `volume_immutable`.
volume_size_gbNohow big that volume may get. Fixed and repeated with `volume_path`, above.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though annotations already declare readOnlyHint=false, openWorldHint=true, and idempotentHint=true, the description adds substantial behavioral context: env is replaced wholesale, volumes are immutable and must be restated, domain changes are refused, and misusing the tool can pull secrets into the conversation. These are exactly the non-obvious behaviors that annotations cannot capture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but structured into a first paragraph on purpose and usage, then a bolded set of rules that directly affect call shape. Every warning has operational value, and the key points are front-loaded. It is not concise in raw length, but it earns its length through dense, decision-relevant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating deployment tool with 10 parameters, no output schema, and many adjacent sibling tools, the description covers the essential preconditions, refusal modes, immutable fields, and alternative tools. An agent has enough context to decide when to call it and how to shape the call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all parameters at 100% coverage, so the baseline is 3. The description adds meaningful beyond-schema semantics for `env` (replaced wholesale, restate every variable), `volume_path` and `volume_size_gb` (immutable and must match `status`), and `deps` (additive, never removes). This justifies a score above baseline, though not the maximum because the per-parameter schema descriptions already carry much of the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action — deploying an image that already exists in gagarin's registry — and clearly distinguishes it from sibling tools like `run`, `add_domain`, and `rotate_resource`. It moves beyond the title by explaining what a deploy actually does and what it refuses to do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: deploy an existing image or a restatement; use `run` for short-lived jobs; use `add_domain` for addresses; use `rotate_resource` for changing a single external key. It also warns against reconstructing the environment from `history`, naming an exclusion and a safer alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

depsWhat a service may reachA
Read-only
Inspect

Its outgoing edges, what depends on it, and how the graph got that way. A private service is default-denied: until the caller declares the edge, its calls are dropped, which hangs rather than failing fast — so this is the first thing to read when something times out talking to something else in the same project.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
serviceYesservice name, unique within the project

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint and openWorldHint. The description adds meaningful behavioral context beyond those: private services default-deny, undeclared calls are dropped, and the failure mode is a hang rather than a fast failure. This helps the agent reason about timeout symptoms without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first frames what the tool shows, the second delivers the critical operational warning. The timeout scenario is placed after the definition, and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only dependency tool with two well-schema'd parameters and no output schema, the description covers what data to expect and why it matters in a real troubleshooting flow. It doesn't spell out the output format, but the content list is sufficient for selecting and invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both project and service are already documented. The description does not add parameter-specific syntax or examples, but it doesn't need to; baseline 3 applies when the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description concretely lists what the tool provides: outgoing edges, dependents, and graph history, so an agent can tell this is a read-only dependency inspection tool. It doesn't start with an explicit verb like 'show' or 'list,' and it doesn't name the sibling set_deps, but the content is specific and not tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear diagnostic trigger: read this first when a service times out talking to another service in the same project, because private services are default-denied until the caller declares the edge. It lacks explicit exclusions or alternatives, but the when-to-use context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

destroy_projectDelete a whole project (needs a human)A
Destructive
Inspect

Everything in it: every service, every resource, every volume, every backup. Answers approval_required and emails the account owner a button; call it again after they have clicked. Say plainly what will be lost before you ask for this.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds crucial behavioral details beyond that: the approval flow (answers approval_required, emails a button, requires a follow-up call) and the caution to clearly communicate losses. This is exactly the kind of non-obvious behavior an agent needs to know, and it is not present in the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences. The first sentence front-loads the destructive scope, the second explains the approval mechanism and retry, and the third gives a direct user-facing instruction. Every sentence earns its place with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with a single parameter, the description covers the full scope of destruction, the approval flow including the retry step, and a caution about user communication. No output schema exists, so no return-value details are needed. An agent has everything necessary to call this correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema fully documents the single parameter (project name or id, 100% coverage). The description adds no extra parameter-level semantics, but none are needed since the parameter is self-explanatory. Baseline 3 is appropriate when the schema covers everything.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb ('Delete') and a specific resource ('whole project') and spells out the full scope ('every service, every resource, every volume, every backup'). It clearly distinguishes itself from sibling tools like destroy_service or destroy_resource by focusing on the entire project, so an agent can differentiate it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use it (when the whole project must be deleted) and provides critical usage context: it returns approval_required, emails the account owner, and requires a second call after the owner clicks. It also advises stating what will be lost before asking. It doesn't explicitly name alternatives, but the scope makes the intended use obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

destroy_resourceDelete a resource and its data (needs a human)A
Destructive
Inspect

Answers approval_required and emails the account owner a button. Call it again after they have clicked. The data goes with it — take a backup first if there is any doubt, and note that services still declaring it will start failing to connect.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
resourceYesresource name, unique within the project

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotation destructiveHint=true, the description explicitly states that the data is gone, recommends a backup, and warns that dependent services will fail to connect. This adds critical context about the irreversible consequences and side effects, which annotations alone do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loading the primary behavior and approval process. Each sentence is purposeful: the first explains the workflow, the second emphasizes destructive consequences, and the third mentions side effects. There is no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, fully documented schema with strong annotations and no output schema, this description is complete. It covers the approval flow, the need to call again, the risk of data loss, and the impact on dependent services. An agent has all necessary information to invoke and anticipate outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of parameter descriptions, so the baseline is 3. The description does not add new parameter semantics, but it is not required given complete schema coverage. No additional meaning is needed beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool destroys a resource and its data, requiring a human approval process. This distinguishes it from sibling tools like destroy_service and destroy_project, which target different resource types. The title reinforces the destructive nature, and the description explicitly notes the action is irreversible.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use it: it answers a pending approval and should be called again after the human clicks the approval button. It indirectly distinguishes itself from destroy_service and destroy_project by naming them as alternatives for other resource types, and advises a backup before use if there is any doubt.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

destroy_serviceDelete a service (needs a human)A
Destructive
Inspect

Answers approval_required and emails the account owner a button. Call it again after they have clicked, within the window the hint names. Deletes the service, its ingress and its volume; the images stay in the registry. Jobs go the same way, taking their kept runs with them.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
serviceYesservice name, unique within the project

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations, which already mark this as destructive. It discloses exactly what gets destroyed (service, ingress, volume), what survives (images), and how jobs and their kept runs are affected. It also explains the two-phase approval behavior, adding substantial behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each packed with essential information: the approval sequence, the deletion scope, and the fate of jobs. No filler or repetition; the most critical operational detail (approval and second call) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with two simple parameters and no output schema, the description covers the important operational steps and side effects. The only notable ambiguity is 'within the window the hint names' — it assumes the agent knows what 'the hint' refers to, which is slightly underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters already have clear descriptions ('project name or id', 'service name, unique within the project'). The tool description adds no parameter-specific meaning beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Deletes'), the resource ('service'), and the exact scope ('its ingress and its volume; the images stay'). It distinguishes itself from sibling tools like destroy_project and destroy_resource by targeting services specifically and by describing service-specific side effects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear procedural context: call it once to trigger approval, then call it again after the account owner clicks, within the named window. It does not explicitly name alternatives or when-not-to-use conditions, but the approval flow and service-specific scope provide enough orientation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ejectTake everything and leaveA
Read-only
Inspect

The Kubernetes manifests, Dockerfiles and connection details for a whole project, so it can be run somewhere else. Owner only: what comes back includes every service's environment in the clear.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds security-relevant context beyond the readOnlyHint annotation: it warns that only the owner can use it and that the returned data includes every service's environment in cleartext. These are important behavioral disclosures not present in the structured annotations, and there is no contradiction with the readOnlyHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and information-dense, with both sentences earning their place: the payload and the security warning. However, the first sentence is a grammatical fragment rather than a complete sentence, which slightly reduces directness and clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers what is returned, the purpose, and the security restriction. It doesn't specify the output format (archive, directory, etc.), but that is a minor omission given the simple interface and the annotations already providing the read-only safety profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'project' parameter is fully described in the schema as 'project name or id', and the description only references 'a whole project' without adding format, examples, or constraints beyond the schema. With 100% schema description coverage, the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the concrete resource (Kubernetes manifests, Dockerfiles, connection details) and the purpose (run elsewhere), but it lacks an explicit verb like 'exports' or 'returns,' relying on the title and 'what comes back' to imply the action. It doesn't explicitly distinguish from siblings like deploy or run, though the resource list is distinctive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for migrating a project by saying 'so it can be run somewhere else' and restricts to owner, but it does not name alternative tools or state when not to use it. There is no explicit guidance about choosing this over siblings like deploy or transfer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

historyEvery deploy of a serviceA
Read-only
Inspect

Each recorded revision with its image, port, environment and the dependencies it ran under. revision is what rollback takes — and for a job it is also what each run was called, so this is the list of runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
serviceYesservice name, unique within the project

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already provide readOnlyHint=true, and the description adds meaningful behavioral detail: each entry includes image, port, environment, and dependencies, and the revision identifier doubles as the run identifier for jobs. There is no contradiction with the read-only annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence front-loads what each revision contains, and the second ties the output to rollback and job runs, making every clause informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains the return contents and the special meaning of `revision` for jobs. It does not mention ordering or pagination, but for a read-only list of revisions with only two required parameters, the description covers what the agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the two parameters `project` and `service` are already fully described there. The description does not add extra semantics about these parameters or their expected formats, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource as recorded revisions with image, port, environment, and dependencies, and explicitly connects `revision` to what `rollback` takes. It lacks an explicit verb like 'list' or 'get', but the title 'Every deploy of a service' and the phrasing 'this is the list of runs' make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage context by stating that `revision` is what `rollback` consumes and that for jobs this tool is the list of runs. It does not explicitly contrast it with `logs` or other siblings, but the rollback and run relationships make when-to-use reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

logsRecent logsA
Read-only
Inspect

The last 200 lines from a service, or from a job's latest run. A tail, not a stream — there is no more, and for a job there is no way to read a run older than the last one.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
serviceYesservice name, unique within the project

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true, so the read-only nature is covered. The description adds valuable behavioral context beyond annotations: it is a tail, not a stream, and there is no way to read a run older than the last one for jobs. This enriches the agent's understanding of limitations without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the core purpose ('last 200 lines') and immediately clarifies limitations ('a tail, not a stream' and no older runs). Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with only two parameters and no output schema, but the description introduces ambiguity: it mentions 'job's latest run' yet the schema only accepts project and service. It does not explain how to specify a job, nor does it describe the output format (e.g., array of lines). These gaps make it incomplete for an agent deciding how to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are fully described in the schema (100% coverage), so the description adds no additional parameter meaning. The description mentions 'service' and 'job' but does not clarify how job is specified given the schema only has project and service. Baseline of 3 applies since schema covers the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves the last 200 lines from a service or a job's latest run. It uses a specific verb ('tail') and distinguishes itself from a stream, making its purpose unambiguous. The distinction from siblings like 'history' is implicit but sufficient.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context that it is a finite tail and that older runs cannot be accessed, but it does not explicitly state when to use this tool versus alternatives like 'history' or when not to use it. It implies usage for recent logs only, but lacks explicit routing to other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

membersWho can reach a projectA
Read-only
Inspect

The owner and everyone it has been shared with, plus a pending offer of ownership when there is one. The owner is a property of the project rather than a row in the list: one account pays, and that is not a role granted or revoked — it moves only through transfer.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds valuable behavioral context: the owner is not a row in the list but a project property, a pending ownership offer is included, and ownership only moves via transfer. This goes beyond what annotations alone convey and helps the agent interpret results correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, each earning its place. The first sentence states the complete result set in a compact way; the second clarifies the unusual owner semantics. No fluff or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read-only tool with no output schema, the description sufficiently conveys the membership model and the special owner case. It stops short of explicitly saying 'returns a list of members' and does not describe the output representation, but the title and text cover the essential semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'project' is fully described in the schema with coverage at 100%. The tool description adds no extra parameter-level information, so the schema carries the burden. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title 'Who can reach a project' and the description together make clear this tool returns the set of accounts (owner, shared users, pending offer) that can access a project. However, there is no explicit verb like 'list' or 'get', so the action is implied rather than stated. It differentiates itself from share/transfer siblings by describing the resulting membership state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the tool's domain model but never explicitly states when to use this tool versus alternatives like share, transfer, or project list tools. Usage is implied by the title and the read-only nature, but there are no when/when-not instructions or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

platform_healthIs gagarin upA
Read-only
Inspect

The platform's own readiness, unauthenticated. Answers whether the control plane can reach its database and cluster and when the reconciler last ran — not whether your service is up, which is status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only and open-world, and the description adds meaningful behavioral context beyond them: no authentication required, what components are checked (database and cluster), and that reconciler timing is included. It could further describe the response shape, but the disclosed semantics are strong for a no-parameter health check.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the core purpose, and every clause earns its place. The exclusion of `status` is compactly integrated rather than bolted on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only health check this is nearly complete: the agent knows what the tool checks, that it is unauthenticated, and how it differs from `status`. The only minor gap is the absence of return-format details, but no output schema exists and the described semantics are sufficient for correct selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and schema coverage is 100%, so the schema carries no burden. The baseline for parameterless tools is 4, and the description appropriately focuses on semantics rather than inventing parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific, informative framing: 'The platform's own readiness, unauthenticated' and explicitly states what it answers (control plane can reach database and cluster, reconciler last ran). It clearly distinguishes itself from the sibling `status` tool, which checks service availability.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states the exact scope of use — platform readiness, not service health — and names the alternative tool (`status`) for the excluded case. The 'unauthenticated' qualifier also tells the agent no credentials are needed, which is direct usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

projectsEvery project you can reachA
Read-only
Inspect

Projects this account owns or has been shared with, and the role on each. A viewer role means every deploy will be refused, which is worth knowing before the attempt. Names are unique only within one account, so two rows can share a name — the id tells them apart, and every other tool takes either.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and openWorldHint. The description adds valuable context about viewer roles blocking deploys and the uniqueness of names within an account, which are behavioral traits beyond the annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose, and every sentence adds value. Warning about viewer role is placed strategically to inform deploy behavior. No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with no output schema, the description covers what is returned (projects and roles), the critical role side-effect, and the id/name usage convention. Slightly more detail about ordering or fields could be added, but it is functionally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has zero parameters, so baseline is 4 per instructions. The description explains what the response contains (projects with roles) and how to use ids/names, adding meaning beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states the tool lists projects the account owns or has shared access to, along with the role on each. It distinguishes itself from siblings like create_project and share by focusing on enumeration, and clarifies the id/name distinction without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description implies usage (e.g., checking roles before deploys) but does not explicitly contrast with sibling tools or specify when to choose this over alternatives like whoami. It provides context but lacks direct 'use this when' guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_domainTake an address awayA
Destructive
Inspect

With a domain, releases that custom name. With none, takes the service off the internet entirely — which is refused while a custom name still points at it, since that would leave somebody's DNS aimed at a host gagarin no longer serves. Needs a human's approval, like every other release.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNothe custom name to release. Absent means the generated address.
projectYesproject name or id
serviceYesservice name, unique within the project

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations: it explains the refusal condition when a custom name points at the host, and notes the human approval requirement. This complements the destructiveHint annotation without contradicting it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences that front-load the core action, then explain the refusal condition and approval requirement. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the two modes, the refusal condition, and approval requirement. It does not describe return values or error cases, but with no output schema and annotations present, this is adequate for a removal operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides descriptions for all parameters, and the description aligns with the schema's note about 'domain' being optional. It adds no new semantic detail beyond what the schema already states, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's dual purpose: releasing a custom domain or removing the service's address entirely, and it distinguishes from sibling tools like destroy_service by clarifying the service itself is not destroyed. It also notes the refusal condition, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each mode (with/without domain) and mentions the refusal condition that requires removing the custom name first. It does not explicitly name alternatives, but the context is clear enough for an agent to choose correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resource_keysWhat a resource publishes, by nameA
Read-only
Inspect

The variable names a resource publishes to whatever declares it — and no values. Prefer this over resource_secrets unless you actually need a value. It answers the question you usually have: what is in this bundle, and what is a key called, before you change one of them with rotate_resource — an external might publish API_KEY or TOKEN and only this says which. The values are never fetched, so nothing secret enters this conversation; status will not tell you either, since it reports only that the resource publishes NAME_*.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
resourceYesresource name, unique within the project

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description discloses that values are never fetched, nothing secret enters the conversation, and status only reports NAME_* patterns. This adds meaningful non-obvious behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence, and subsequent sentences add routing and behavioral caveats. It is slightly discursive but every clause contributes useful context; there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only name-listing tool with fully documented parameters and strong annotations, the description tells the agent what to expect (names, no values), when to use it, and what it will not do. Even without an output schema, it is complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are simple, well-described identifiers. The description adds domain context but no parameter-specific format, constraints, or additional semantics, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific outcome: the variable names a resource publishes, with no values. It also distinguishes itself from resource_secrets and relates to rotate_resource, making the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'Prefer this over resource_secrets unless you actually need a value', giving a clear when/when-not rule. It also explains why this is useful before rotate_resource and why status is insufficient, so an agent can route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resource_secretsA resource's connection valuesA
Read-only
Inspect

The host, port, user, password and URL a client would need. For reading, and for something outside the project: a service inside it should declare the resource in deps instead, which hands the same values over without anybody copying a password. This returns live credentials into this conversation, so call it only when a value is what you need — to hand a connection string to something outside the project, or to verify a rotation. If you only need to know what the resource publishes, or what a key is called, that is resource_keys and it returns no values.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id
resourceYesresource name, unique within the project

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, but the description adds crucial context beyond that: it returns live credentials into the conversation and carries a warning to call only when a value is truly needed. It also explains that the deps alternative avoids password copying, which is a meaningful behavioral nuance. It does not cover side effects like logging or audit trails, but the core sensitive behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense, with the returned values front-loaded, the internal-vs-external use case explained, and a clear warning in bold. Every sentence earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema, the description adequately explains what the caller receives (host, port, user, password, URL), when to use it, when to avoid it, and which sibling to choose instead. The annotations cover read-only behavior, and the warning addresses the sensitive nature. An agent has enough context to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters (project, resource) have descriptions in the schema itself. The tool description does not add new semantic meaning for these parameters beyond what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact resource ('host, port, user, password and URL') and identifies the operation as returning live credentials. It clearly differentiates itself from resource_keys, which returns no values, so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to call this tool (when a literal value is needed, e.g., handing a connection string outside the project or verifying a rotation) and when not to (services inside the project should use deps; metadata-only needs should use resource_keys). This is concrete, actionable routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_resourceRestore a backup into a new resourceAInspect

Fills a new resource from another one's backup and never overwrites anything — which is why it needs no approval and cannot lose data. Three calls, in this order, and the first two are not optional: add_resource to create the destination with the same type as the source, status until it is running, then this. A restore into a name that does not exist answers no_such_resource, and one into a resource that exists but has not come up yet answers restore_failed — the data is loaded into a live resource, not conjured as one. resource is that new name; source is the resource whose newest backup to take, or backup is one exact key from backups. Point the dependents at the new name with set_deps once you have checked it.

ParametersJSON Schema
NameRequiredDescriptionDefault
backupNoone exact key from `backups`, for a specific point in time
sourceNothe resource whose newest backup to use
projectYesproject name or id
resourceYesthe new resource to fill. Create it with `add_resource` and wait for `status` to show it running before calling this.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as not read-only and not idempotent, but the description adds substantial behavioral context: no approval required, no data loss possible, data is loaded into a live resource, and specific error responses (`no_such_resource`, `restore_failed`). There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: it covers the core guarantee, the required call sequence, failure modes, parameter roles, and the follow-up `set_deps` step. The use of backticks and the three-step structure makes it easy to parse despite the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex mutating workflow with no output schema, but the description covers prerequisites, ordering, error semantics, parameter selection, and post-steps. An agent has enough context to invoke this tool correctly and to react to expected failures.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, which sets a baseline of 3, but the description adds meaningful semantics beyond the schema: `resource` is the new destination, `source` selects the newest backup, and `backup` selects an exact key from `backups`. This helps disambiguate two optional-looking parameters, though it stops short of explicitly stating that exactly one of `source` or `backup` must be supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('restore a backup into a new resource') and immediately distinguishes this from overwriting behavior: it 'never overwrites anything' and fills a brand-new resource. This differentiates it from sibling tools like `rollback`, `backups`, and `destroy_resource` without needing to inspect those schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit orchestration context: `add_resource`, then `status` until running, then this tool. It also explains expected failure responses for invalid usage. It does not explicitly name an alternative tool for in-place restore scenarios, but the 'never overwrites' framing clearly implies this tool is for the non-destructive path, so the guidance is nearly complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_credentialTake a credential awayA
Idempotent
Inspect

Stops a credential working immediately. credentials gives the id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesthe id from `credentials`

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and idempotentHint=true, so the description's mutation effect aligns. It adds the nuance of 'immediately' and the outcome of stopping working, which is useful. However, it doesn't disclose potential side effects like invalidating sessions or reversibility. With annotations covering the basics, the extra context is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste. The core action is front-loaded, and the id source is given as a follow-up hint. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter revoke tool with no output schema, the description covers the essential: what it does and where to get the id. It doesn't describe edge cases like already-revoked credentials, but given the low complexity and supporting annotations, it is sufficiently complete for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already states 'the id from `credentials`'. The tool description repeats this same hint, adding no new information. Baseline for high coverage is 3, and there's no extra semantic detail provided beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Stops a credential working immediately') and clearly identifies the resource (credential). It distinguishes itself from siblings like create_credential and rotate_resource by focusing on revocation. The verb 'Stops' is precise and the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use (when you need to stop a credential) and provides the source of the id via `credentials`. It doesn't explicitly mention alternatives or exclusions, but for a simple revoke action the context is clear enough. There is no misleading guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rollbackPut a past revision backAInspect

Deploys a revision this service already ran. No human approval, because it restores a state that was already approved once; it refuses to cross a change of volume. Rolling a job back is not a restoration but a re-run of that earlier image, under a new revision — so only ask for one if running it twice is safe. It restores the image and the environment of the deploy, and NOT the variables the service inherits from resources it needs — those are resolved from the graph as it stands now, so a rollback never puts a service back onto a rotated password. To undo a config change, roll back the external resource holding it, not its dependents: name the resource here and every service declaring it is restarted with the restored values. An external can be rolled back because its values are the user's; a postgres, qdrant or valkey cannot, because gagarin mints those and there is no earlier value of theirs to return to.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNothe revision from `history`. Absent means the previous one.
projectYesproject name or id
serviceYesthe service to roll back — or the name of an external resource, to put its previous values back across everything that declares it.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, idempotentHint=false), the description discloses crucial behavioral nuances: it restores image and environment but NOT inherited variables, refuses to cross a volume change, and re-runs under a new revision rather than restoring state. It also explains the external-resource behavior and why managed resources cannot be rolled back. This adds substantial transparency beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly packed with essential information, front-loading the core purpose and then elaborating on edge cases. Bolded phrases highlight key distinctions. No sentence is wasted; the length is justified by the tool's complexity. Slightly less verbose could be possible, but the structure serves the content well.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description covers all critical aspects an agent needs: what is restored, what is not, when to use it, how to handle external resources, and which resources are ineligible. It addresses potential pitfalls like rotated passwords and volume changes. Nothing essential appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the input schema provides 100% coverage for all three parameters, the description enriches their meaning significantly. It clarifies that 'service' can be an external resource name to roll back its previous values across dependents, and it explains the 'to' parameter's default behavior (previous revision) in context. This goes well beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Deploys a revision this service already ran') that precisely states the action. It also distinguishes itself from related operations by clarifying that a job rollback is a re-run rather than a restoration, and by explaining how to undo config changes via external resources. This differentiates it from siblings like restore_resource and rotate_resource without needing to name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: it explains that to undo a config change one should roll back the external resource, not its dependents, and it lists which resources are eligible (externals) versus ineligible (postgres, qdrant, valkey). It also warns that a job rollback is safe only if running the earlier image twice is safe, providing clear conditions and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rotate_resourceReplace a resource's credentialsAInspect

New credentials, and everything holding them rolls to pick them up. For an external the new values are yours to supply and required; for everything else they are gagarin's to mint and supplying them is refused. An external usually holds several values, and there are two ways to change them: set/unset change the keys you name and leave the rest exactly as they are, while env says the bundle is now precisely this and drops every key not in it. Reach for set when one key is being replaced — env with a single key would take the others away from every dependent. The answer names what changed and what stopped being published, so read removed before reporting success.

ParametersJSON Schema
NameRequiredDescriptionDefault
envNoeverything the resource should publish from now on, replacing what is there. Required for an `external` unless set/unset is given, refused for every other type. Anything omitted stops being published.
setNochange these values and leave every other key alone. What to use when one secret of several is being rotated. `external` only, and not combinable with env.
unsetNostop publishing these keys, named without the resource prefix. Refused if the resource does not publish one of them, so a mistyped name fails rather than silently doing nothing. `external` only, and not combinable with env.
projectYesproject name or id
resourceYesresource name, unique within the project

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is a mutating, non-idempotent operation, and the description adds significant behavioral detail: the propagation of new credentials, the refusal of user-supplied values for non-external types, the key-preserving vs. key-dropping behavior of set/unset vs. env, and the presence of a 'removed' field in the response. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: it opens with the core behavior, then explains the external vs. non-external distinction, then details the three modes, and ends with a pointer to the response's `removed` field. The bolded recommendation for `set` is front-loaded and effective. A bit long, but each sentence carries necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of five parameters with nested objects and the conditional logic between resource types, the description covers all key decision points: when credentials are required vs. refused, how each mode behaves, and what to read in the response. No critical information is missing for an agent to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are documented, but the description adds semantic nuance beyond the schema—explaining how `env` replaces the entire bundle, how `set`/`unset` preserve other keys, and the exclusivity between `env` and `set`/`unset`. This helps an agent choose the right parameter combination.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title and description clearly state the tool replaces a resource's credentials, and the description elaborates on the mechanics of rotation for external vs. non-external resources. It distinguishes the operation from generic create/revoke by focusing on the 'roll' behavior, though it doesn't explicitly name sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear guidance on when to use `set` vs `env` and explains the constraints for `external` resources, but it does not explicitly state when to prefer this tool over create_credential or revoke_credential. The usage context is implied rather than contrasted with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

runRun an image to completionAInspect

Submits a job: an image that runs, exits, and is done — a migration, a backfill, a one-off script. Not a service. It has no port and no volume, nothing can be told to reach it, and it is never restarted by the platform: a script that exits non-zero is reported as having failed, once, with its code. Like deploy, the image must already be in gagarin's registry and this server cannot put it there. And like deploy, this returns without waiting. It answers a revision, which is the run's name; status reports that run's phase and exit code, and logs reads it. There is no tool here that blocks until a run finishes — gg run does that on a machine, exiting with the script's own code, which is what CI wants. Calling this again is the next run, not a restatement: each one is a new revision, and the three most recent are kept so the previous run's logs survive. A name that is already a service is refused not_a_job — the two are not two states of one thing.

ParametersJSON Schema
NameRequiredDescriptionDefault
envNothe complete environment for this run. Absent empties it — not merged.
depsNoresources and services this run may reach while it runs, added to whatever it already declares. A migration needs its database named here, or its connection hangs rather than failing.
sizeNoCPU and memory envelope. A job is billed for the time it runs, at this size.
imageYesfull reference in gagarin's registry, e.g. registry.gagarin.cloud/<project-id>/migrate:v3
digestNosha256:... as docker push reported it, pinning the exact image
projectYesproject name or id
serviceYesjob name, unique within the project among services too

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false, openWorldHint=true, idempotentHint=false. The description goes far beyond: it discloses non-blocking return, no ports or volumes, no restart, one-time failure reporting with exit code, a new revision per call, three-run retention, and the not_a_job refusal. This gives the agent a precise behavioral model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although long, every sentence carries a distinct fact: job semantics, deploy similarities, non-blocking behavior, revision lifecycle, and naming conflicts. The core purpose is front-loaded in the first sentence, and bold key terms aid scanning. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still explains that the tool returns a revision and points to status and logs for follow-up. It also covers failure reporting, non-idempotency, retention, and the name-conflict error, making the tool fully understandable and callable from the text alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter already has a meaningful description. The description adds the prerequisite that the image must already be in the registry and that service names are refused, but these are largely implied by the schema's 'full reference in gagarin's registry' and 'unique within the project among services too.' The added parameter-level insight is marginal, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Submits a job: an image that runs, exits, and is done' and immediately distinguishes it from a service ('Not a service'). It lists concrete use cases (migration, backfill, one-off script) and references the sibling deploy to frame what it is not, making it easily distinguishable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states this is not a service and that service-named jobs are refused, and it references deploy for shared constraints. However, it never explicitly instructs 'use deploy for services' or names the exact alternative for long-running workloads, and the blocking alternative (gg run) is an external CLI rather than a sibling tool. The guidance is strong but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_depsDeclare what a service may reachA
Idempotent
Inspect

Replaces the complete set of things this service may reach, so an empty list means "nothing" and is a real request. Declaring a resource opens the route and hands over its connection variables, so connecting a database is this one call and not a call plus a deploy — the dependents roll on their own. To add without risking a withdrawal, pass deps on deploy instead. A set that drops an edge this service currently holds needs a human: it answers approval_required and emails the owner, exactly as deleting a service does. A set that only adds goes through. Read deps first and send that list plus your additions, or you will ask somebody to approve a withdrawal you did not mean — and a withdrawal is the one change that reports nothing at runtime: the calls are dropped, not refused, so the far end hangs until it times out. A job may need things, and nothing may need a job: an edge is a rule about a port and a job has none, so naming one here is refused. If a job and a service share data, that is a resource they both need.

ParametersJSON Schema
NameRequiredDescriptionDefault
needsYesthe complete list of services and resources this one may reach
projectYesproject name or id
serviceYesservice name, unique within the project

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover readOnly, openWorld, and idempotency, yet the description goes far beyond them: it discloses the side effect of handing over connection variables, the human-approval gate with `approval_required` and owner email, the silent failure mode (calls dropped, not refused, so the far end hangs), and the job/edge domain restriction. All disclosed behaviors are consistent with the annotations — replacement semantics align with idempotentHint, and readOnlyHint=false matches the mutating nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense paragraphs, each topically organized: replacement semantics and alternatives, approval/safety workflow, and domain restrictions. Every paragraph earns its place given the tool's subtlety, and the core replacement semantics are front-loaded. Some phrasing is indirect ("the dependents roll on their own") and the withdrawal warning is repeated, so it is not maximally tight — but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and this much semantic subtlety, the description covers the critical ground: replacement behavior, approval gating, the alternative path, the silent-failure mode, and the job restriction. The main gap is the success response — the description mentions `approval_required` but never states what a successful, non-approval response returns or how the agent confirms completion.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds semantics the schema cannot express: `needs` is a full replacement (not an append), an empty list is a real request meaning "nothing", naming a job in `needs` is refused, and declaring a resource also imports its connection variables. This materially changes how an agent must construct the `needs` parameter, going well beyond the schema's "complete list" phrasing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: "Replaces the complete set of things this service may reach." It is clearly distinct from the sibling `deps` (the read tool) and from `deploy` (which takes `deps` additively), and the title "Declare what a service may reach" reinforces the purpose. An agent can tell this tool from its siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative and the selection condition: "To add without risking a withdrawal, pass `deps` on `deploy` instead." It also prescribes the safe workflow — "Read `deps` first and send that list plus your additions" — and explains the consequence of ignoring it (triggering an unintended human-approved withdrawal). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shareGive somebody accessA
Idempotent
Inspect

Grants a role on a project, or changes one somebody already has. An editor can do everything the owner can except be billed for it; a viewer reads.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYeseditor or viewer
emailYestheir email address
projectYesproject name or id

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal mutation (readOnlyHint=false), idempotency, and open-world side effects, so the bar is lower. The description adds meaningful behavioral context by defining what an editor can do versus a viewer, and by noting that the tool can either grant or alter an existing role. It does not contradict any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler. The primary action is front-loaded, and the role distinction is explained precisely in the second sentence without expanding the definition unnecessarily.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter tool with no output schema, the description covers the core operation, the update case, and the meaning of the role parameter. It could be slightly more complete by referencing removal or ownership-transfer tools, but nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 even though the description adds no new parameter-level detail. The role semantics in the description reinforce the schema's 'editor or viewer' enum but do not go beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Grants'), a clear resource ('a project'), and the object of the action ('a role' on behalf of somebody), and it additionally covers the update case ('or changes one somebody already has'). This distinguishes it from nearby siblings like unshare (revoking) and transfer (ownership handoff), so an agent can tell it apart without opening another tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for the share update behavior ('or changes one somebody already has') and explains role semantics, which helps an agent decide between viewer and editor. However, it never explicitly says when to prefer this over siblings like unshare, transfer, or members, leaving alternative selection mostly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusDesired versus actualA
Read-only
Inspect

The only call that reads the cluster, and so the only one that can answer "is it up". Every write on this server is asynchronous — a tool that returns without error recorded a demand, it did not watch it come true — so this is what you check afterwards. Carries every service and resource, its addresses, its size, whether it is in sync, and what the project has cost since midnight UTC. A job has none of the service vocabulary — no ready count, no port, no address — and carries its latest run instead: which revision, what phase, how long it took and the exit code. That is the only way to learn how a run ended. It does not carry a resource's environment: what a resource publishes by name is resource_keys, and the values are resource_secrets.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, it discloses key behavioral context: all writes are asynchronous, return success only records a demand, and this endpoint is the post-write verification point. It also explains the differences between service and job payloads and explicitly notes what the response does not include.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence adds essential behavioral or semantic information. It front-loads the core purpose and then systematically covers write-asynchrony, response contents, job differences, and exclusions without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description compensates well by explaining what the tool returns for services, resources, and jobs, as well as what it intentionally omits. This is sufficient for an agent to correctly select and interpret the tool in most workflows.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes the single 'project' parameter with 100% coverage. The description does not add extra meaning about the parameter itself, but no compensation is needed because the schema is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads the cluster and is the only call that can answer 'is it up'. It distinguishes itself from sibling tools by emphasizing its unique read-only role and by explaining what data it carries for services, resources, and jobs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: it is the tool to check after asynchronous writes, the only way to verify if the cluster is up, and the only way to learn how a run ended. It also implicitly tells agents when not to use this tool by noting it does not carry resource environments, which belong to resource_keys and resource_secrets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transferOffer a project, and its bill, to somebodyA
Idempotent
Inspect

Offers ownership to somebody the project is already shared with. This does not hand it over: it emails them, and the project moves only when they press the button, which may be days later or never. Two humans are involved — the owner is asked to approve the offer in their own inbox first (approval_required), and the recipient accepts in theirs. When it lands, the previous owner stays on as an editor and nothing restarts. Never call this unless the user has asked for the project to change hands.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNowhat to call it in their account, if they already have one by this name
emailYestheir email address; they must already be a member
projectYesproject name or id

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral detail beyond the annotations: it emails the recipient, requires two human approvals, may complete days later or never, and leaves the previous owner as an editor. This fully discloses the asynchronous and conditional nature of the mutation, with no contradiction to the given hints (readOnlyHint=false, openWorldHint=true, idempotentHint=true).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: it front-loads the core purpose, then explains the two-step approval flow, post-transfer state, and adds a strong usage warning. Slightly verbose but never redundant, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description covers the complete workflow: what triggers it, the async email flow, the need for two approvals, the eventual ownership change, and the final role of the previous owner. It also gives an unambiguous invocation condition ('Never call this unless...'), so an agent has all context needed to decide and call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter (name, email, project) already has a clear description in the schema. The tool description adds no extra parameter-level meaning; it only restates that the recipient must already be a member, which is already in the email parameter description. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Offers ownership to somebody the project is already shared with.' It also differentiates from siblings by clarifying 'This does not hand it over' and noting the recipient must already be a member, which distinguishes it from share or immediate transfer tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance is provided: 'Never call this unless the user has asked for the project to change hands.' It also states the prerequisite that the recipient must already be shared with the project, giving clear when-to-use conditions and implicitly distinguishing from share/untransfer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unshareTake access awayA
Idempotent
Inspect

Removes somebody from a project. The owner cannot be removed.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailYestheir email address
projectYesproject name or id

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a mutating operation (readOnlyHint=false) and idempotent (idempotentHint=true). The description adds one meaningful behavioral fact—the owner cannot be removed—which goes beyond the annotations. However, it does not disclose outcomes for invalid emails, whether removal is reversible, or any side effects, leaving room for more transparency given the openWorldHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loading the core action and then adding the critical constraint. Every word earns its place, and there is no redundant filler. This is appropriately concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation tool, the description covers the essential behavior (removal), the required inputs, and an important edge case (owner). It does not describe return values, but with no output schema and strong annotations, this is not critical. A more complete version could mention error handling or reversibility, but the definition is largely sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: both 'project' and 'email' have descriptions ('project name or id', 'their email address'). The description does not add additional semantic detail beyond what the schema already provides, so it meets the baseline of 3 but does not exceed it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action: 'Removes somebody from a project', with a specific resource (project) and subject (somebody). It also includes a key constraint ('The owner cannot be removed') that helps distinguish this from transfer or share tools. However, it does not explicitly name sibling tools or contrast its scope beyond the owner restriction, so it slightly misses the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this tool to remove a person from a project. The owner restriction provides a conditional, but there is no explicit guidance on when to prefer alternatives like transfer (for ownership changes) or revoke_credential (for credential removal). The description relies on inference rather than offering clear when-to-use versus when-not-to-use direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

untransferTake back an offer of ownershipA
Idempotent
Inspect

Withdraws an offer that has not been accepted, and kills the link in the recipient's inbox. There is no undoing one that has been accepted: the project is theirs, and only they can offer it back.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesproject name or id

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly=false and idempotent annotations, the description discloses a concrete side effect: it 'kills the link in the recipient's inbox' and states the irreversibility boundary for accepted offers. It does not discuss permissions or what happens on repeated calls, but the idempotent hint already covers that aspect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences front-load the action and key limitation, with every phrase adding information. There is no filler, and the critical 'not accepted' condition appears immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, no-output-schema tool, the description covers what the call does, when it applies, and an important limitation, while annotations cover idempotency and non-read-only behavior. An agent has enough to invoke it correctly without extra documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the single project parameter with 100% coverage as 'project name or id', and the description adds no additional parameter-level meaning. This meets the baseline but provides no extra semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb plus resource: it withdraws an unaccepted ownership offer tied to a project, and explains the 'untransfer' action by describing the inbox link destruction. This clearly distinguishes it from ownership-transfer and sharing siblings, since it is specifically about taking back an offer, not access sharing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear condition: use it for an offer that has not been accepted, and an explicit when-not: there is no undoing an accepted offer. It stops short of naming the sibling tool (e.g., transfer) as the alternative, though 'only they can offer it back' implies the recipient would need to transfer it back.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiWho this credential isA
Read-only
Inspect

Which gagarin account this server is acting as, what the credential may do, and the registry and base domain to build addresses from. Run this first in any session: it is the one call that distinguishes "no credential" from "a credential that cannot deploy".

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/openWorldHint annotations, the description adds meaningful behavioral context: the call is safe to run first, reveals the credential's authorization scope, and returns the registry/base domain needed for later address construction. It also implies the call is always executable and distinguishes authentication failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first lists exactly what the call reports, and the second states its privileged position in a session. Important operational guidance is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only identity tool, this is complete: it tells the agent what information it will get, why that information matters for building subsequent calls, and when to invoke it. No output schema exists, but the description covers the return semantics sufficiently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to clarify; the baseline of 4 applies. The description instead explains the value each returned piece of identity information provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool reports—the acting gagarin account, the credential's capabilities, and the registry/base domain to construct addresses. It names the resource and scope, and the identity focus clearly distinguishes it from siblings such as status or platform_health.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: 'Run this first in any session.' It also clarifies the diagnostic value—distinguishing 'no credential' from 'a credential that cannot deploy'—which tells an agent why this call should precede deploy and other operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 35 tool updatesv0.1.0
    • First observedadd_domain
    • First observedadd_resource
    • First observedbackup_resource
    • First observedbackups
    • First observedbilling
    • First observedbilling_history
    • First observedcreate_credential
    • First observedcreate_project
    • First observedcredentials
    • First observeddeploy
    • First observeddeps
    • First observeddestroy_project
    • First observeddestroy_resource
    • First observeddestroy_service
    • First observedeject
    • First observedhistory
    • First observedlogs
    • First observedmembers
    • First observedplatform_health
    • First observedprojects
    • First observedremove_domain
    • First observedresource_keys
    • First observedresource_secrets
    • First observedrestore_resource
    • First observedrevoke_credential
    • First observedrollback
    • First observedrotate_resource
    • First observedrun
    • First observedset_deps
    • First observedshare
    • First observedstatus
    • First observedtransfer
    • First observedunshare
    • First observeduntransfer
    • First observedwhoami

TDQS

A3.9/5.0

Scored across 35 tools

Disambiguation5/5

Each tool targets a distinct resource or action, and the descriptions carefully separate near-neighbors: status vs platform_health, deploy vs run, resource_keys vs resource_secrets, and backups vs backup_resource vs restore_resource. An agent should be able to pick the right tool reliably.

Naming Consistency5/5

Names follow a predictable read/action convention: plural nouns for listing and inspection (projects, credentials, status, backups, members) and verb_object for mutations (create_project, destroy_service, add_domain, set_deps, rotate_resource). The few one-word verbs like deploy, run, and whoami are conventional and do not break the pattern.

Tool Count2/5

35 tools is a large surface for an agent, beyond the 25+ threshold even though the underlying platform is broad. Many tools are narrow variants of the same lifecycle (destroy_service/resource/project, add_domain/remove_domain, resource_keys/resource_secrets), which adds coherence but also bulk.

Completeness4/5

The surface covers the full lifecycle for the apparent domain: projects, sharing and ownership, credentials, deploy and jobs, domains, dependencies, managed resources, backups and restore, and billing. Minor gaps exist—no direct resource resize/retype, no project rename, and no explicit restart other than redeploying—but agents can work around them.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Exposes two MCP tools (discover and execute) that enable agents to query an OpenAPI schema via natural language and execute matched API operations.
    2
    -
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to interact with Scout Live platform capabilities through standardized MCP primitives, including tools for app management, deployment, and logging.
    6
    12 npm
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables an LLM to dynamically discover and call tools across multiple MCP servers (file, GitHub, SQL, Python execution) with authentication, rate limiting, and observability, supporting parallel execution and secure deployment.
    -