Skip to main content
Glama
railyard-sh

Railyard MCP Server

by railyard-sh

Railyard MCP server

A Model Context Protocol server that gives an MCP client (Codex, Claude Desktop, Claude Code, or any other) read and write access to your Railyard projects and organisations. It talks to a running Railyard backend over its REST API and authenticates with a personal access token (PAT).

It speaks MCP over stdio and is written in TypeScript against the official @modelcontextprotocol/sdk.

Just want to install it? Jump to Install below, or follow the standalone INSTALL.md — get a token, paste one config block, verify with whoami.


What it can do

Projects

Tool

Kind

Description

list_projects

read

Projects in an org (id, name, slug, updated-at).

get_project

read

A project's full JSON document and revision ETag, by id or slug.

check_project_name

read

Whether a name is free in an org (and the slug it would get).

create_project

write

Create a new, empty project and save it.

update_project

write · destructive

Conditionally save a project via full-document PUT. Merges partial fields by default; can replace the whole document.

rename_project

write

Change a project's name + URL slug.

delete_project

write · destructive

Permanently delete a project. No undo.

move_project

write

Move a project into another org you can write to, optionally renaming it in the same step.

Validation & export — these operate on a document, so each takes either a saved project (ref) or an inline project you have not saved yet. Neither changes anything stored.

Tool

Kind

Description

validate_project

read

Current server-side rack-layout/site and power findings for a saved or inline project; structural hierarchy/cabling errors are rejected first.

list_export_formats

read

The export targets this build supports (nautobot-csv, netbox-csv, designbuilder-yaml, json).

export_project

read

Render a project into a format and return the files' content, unresolved placements and warnings.

Organisations, members & billing

Tool

Kind

Description

whoami

read

The user your token authenticates as (id, email, name).

server_info

read

Server health: version, schema version, and whether persistence, auth and billing are configured.

list_orgs

read

Organisations you belong to — id, slug, role, plan, billing status.

create_org

write

Create a shared org; you become its owner.

rename_org

write · owner

Change an org's display name.

delete_org

write · destructive · owner

Delete a shared org and every project in it. No undo.

get_org_catalog

read

The org's shared device-type library and revision ETag.

set_org_catalog

write · destructive

Conditionally replace that library wholesale (not a merge).

get_org_roles

read

The org's shared device-role vocabulary and revision ETag.

set_org_roles

write · destructive

Conditionally replace the shared role vocabulary wholesale.

list_members

read

Roster: user id, email, role, joined-at.

set_member_role

write · owner

Change a member's role.

remove_member

write · destructive · owner

Remove a member and drop their live sessions.

list_invites

read · owner

An org's pending invitations.

invite_member

write · owner

Invite an email at a role (needs a current Team/Enterprise plan).

revoke_invite

write · owner

Withdraw a pending invitation.

list_my_invites

read

Invitations addressed to your email.

accept_invite

write

Accept one, joining that org.

get_billing

read

Plan, status, seats, trial/period end, and whether the org is currently entitled to edit.

billing_manage_url

write · owner

Mint a Stripe Checkout or Customer Portal URL to open in a browser. Creates a link only — it charges nothing.

The destructive tools (update_project, delete_project, delete_org, set_org_catalog, set_org_roles, remove_member) are annotated with the MCP destructiveHint, so clients that surface tool safety hints will flag them.

Not exposed, deliberately. Personal-access-token management, account deletion and the starter-example claim are gated to an interactive browser session server-side — a token cannot drive them (see Auth model). The OAuth/magic-link routes and the Stripe webhook are not client-callable. Live collaboration is a WebSocket protocol rather than request/response, so it has no tool; see the caveat on concurrent edits below.

Org selection. Every org-scoped tool accepts an optional org argument (an org id, slug, or name). When omitted it falls back to the RAILYARD_ORG environment variable, and if that too is unset, to your first (personal) organisation. Slugs/names are resolved to the org id the API needs (via GET /api/orgs) automatically.


Related MCP server: Linear MCP Server

Setup

1. Requirements

  • Node.js 20 or newer.

  • A Railyard account and personal access token. The server connects to https://railyard.sh by default. Set RAILYARD_BASE_URL only when using a self-hosted or local backend with persistence and authentication enabled.

2. Mint a personal access token

Quick path:

npx -y railyard-mcp auth

That opens Railyard in your browser. Sign in, open User settings if needed, create a token, and copy the ry_… secret. In headless environments, run npx -y railyard-mcp auth --no-open and copy the printed URL.

Manual path:

  1. Sign in to Railyard in your browser.

  2. Go to User settings → Personal access tokens.

  3. Create a token, give it a name, and copy the ry_… secret — it is shown once, at creation. The server only stores its hash; you cannot retrieve it again.

Treat this secret like a password (see Auth model below).

3. Install and build

Only needed to run from source (or to develop). If you install the published package with npx -y railyard-mcp, skip this; npm fetches the packaged dist/ files for you.

cd railyard-mcp
npm install
npm run build

This compiles src/ to dist/. The entry point is dist/index.js.

4. Configure the environment

Variable

Required

Meaning

RAILYARD_TOKEN

yes

Your ry_… personal access token.

RAILYARD_BASE_URL

no

Railyard base URL. Defaults to https://railyard.sh; override it only for self-hosted or local Railyard.

RAILYARD_ORG

no

Default org (id or slug) for org-scoped tools.

You can smoke-test it from a shell:

RAILYARD_TOKEN=ry_xxx npm start
# (it waits on stdio for an MCP client; Ctrl-C to exit)

Install

Pick your client and paste one config block. For a friendly step-by-step walkthrough see the standalone INSTALL.md; the essentials are below.

Two ways to run it:

  • Published (recommended): npx -y railyard-mcp downloads and runs the package on demand — no clone, no build. Requires the package to be on npm (see For operators if it isn't yet).

  • From source (works today): run the built entry point directly with node /absolute/path/to/railyard-mcp/dist/index.js after npm install && npm run build in this repo (see Setup). Substitute that command/args in any snippet below.

The published package uses the hosted URL https://railyard.sh automatically. For a self-hosted or local backend, add RAILYARD_BASE_URL with your own URL (for example, http://localhost:8080). RAILYARD_ORG is optional — add it to pin a default organisation.

Claude Desktop — one-click bundle (.mcpb)

The easiest path, no JSON. Open Claude Desktop → Settings → Extensions, then drag in (or Install extension) the packaged railyard-mcp.mcpb bundle and fill in the token. Leave the optional base URL at its hosted default unless you self-host. The bundle is built from manifest.json — see For operators.

Claude Desktop — manual config

Add the server under mcpServers in claude_desktop_config.json:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

{
  "mcpServers": {
    "railyard": {
      "command": "npx",
      "args": ["-y", "railyard-mcp"],
      "env": {
        "RAILYARD_TOKEN": "ry_your_token_here",
        "RAILYARD_ORG": "my-team-slug"
      }
    }
  }
}

Restart Claude Desktop after editing. The Railyard tools then appear in the tools menu.

From source: replace the two command lines with "command": "node", "args": ["/absolute/path/to/railyard-mcp/dist/index.js"].

Claude Code

Register it in one command:

claude mcp add railyard \
  --env RAILYARD_TOKEN=ry_your_token_here \
  -- npx -y railyard-mcp

Check it with claude mcp list. Add --scope project to write a shared .mcp.json instead of your user config (keep real tokens out of committed files). From source, swap the trailing -- npx -y railyard-mcp for -- node /absolute/path/to/railyard-mcp/dist/index.js.

A project-level .mcp.json takes the same shape as the Claude Desktop block above.

Codex CLI

Need Codex first? Install the Codex CLI:

curl -fsSL https://chatgpt.com/codex/install.sh | sh

Register it in one command:

codex mcp add railyard \
  --env RAILYARD_TOKEN=ry_your_token_here \
  -- npx -y railyard-mcp

Check it with codex mcp list. Run codex mcp --help to see the rest of the Codex MCP commands. From source, swap the trailing -- npx -y railyard-mcp for -- node /absolute/path/to/railyard-mcp/dist/index.js.

codex mcp login railyard is only for MCP servers that advertise OAuth. Railyard currently uses a PAT, so use npx -y railyard-mcp auth to open Railyard in your browser, then paste the resulting token into the RAILYARD_TOKEN env var above.

Cursor

Edit ~/.cursor/mcp.json (global) or .cursor/mcp.json (project), then enable railyard under Settings → MCP:

{
  "mcpServers": {
    "railyard": {
      "command": "npx",
      "args": ["-y", "railyard-mcp"],
      "env": {
        "RAILYARD_TOKEN": "ry_your_token_here"
      }
    }
  }
}

Any other stdio MCP client

Launch this command with the environment set; the client speaks MCP to it over stdio:

command: npx
args:    ["-y", "railyard-mcp"]
env:     RAILYARD_TOKEN=ry_your_token_here
         RAILYARD_ORG=my-team-slug        # optional

Prefer not to commit real tokens. Keep RAILYARD_TOKEN in a private/user-scoped config, or inject it from your environment rather than checking it into a shared config file.

Verify

Run the whoami tool (or ask "who am I on Railyard?"). It returns your Railyard user id, email and name — confirming the token and URL work. Then try list_projects.


Auth model — why a PAT

A personal access token is the right credential for an MCP server; a session cookie is not.

  • Non-interactive. An MCP server runs headless. It cannot complete an interactive SSO/OAuth or magic-link sign-in to obtain a session cookie, and a copied cookie is a short-lived, browser-bound artefact that expires and can't be rotated cleanly. A PAT is a long-lived credential minted for programmatic use — exactly this case.

  • It's the backend's intended programmatic credential. Railyard's API accepts Authorization: Bearer ry_… on every org-scoped route as a first-class alternative to the browser session cookie. This server sends that header on every request.

  • Safer blast radius by design. Railyard deliberately gates token management itself (creating or revoking PATs) behind an interactive browser session only — a PAT cannot mint or revoke tokens. So even if this server's token leaked, an attacker could not use it to create more tokens or lock you out of revoking it; you revoke it from the browser.

What the token carries. A PAT authenticates as you, across all your organisations, with your full role in each. There are no per-token scopes or expiry yet — so:

  • Treat the token like a password. Don't commit it, log it, or paste it into shared configs. This server never writes the token to its logs.

  • Scope it operationally. Only point this server at orgs you intend it to touch (set RAILYARD_ORG, and be deliberate with write tools). Remember the token can still reach any org you belong to if a tool call names one.

  • Rotate on suspicion. If a token may be exposed, revoke it in User settings → Personal access tokens and mint a new one. Revocation is immediate.

Future hardening (not built yet): per-token scopes (e.g. read-only, or org-restricted) and configurable expiry would let you hand this server a narrower credential. Today a PAT is all-or-nothing, which is why the guidance above matters.


How org access & errors map

  • X-Org-Id header. Project-scoped calls send the resolved org id in X-Org-Id; the org-management routes carry it in the path instead. Either way the backend membership-checks it and returns 403 if the token's user isn't a member.

  • Roles. Reads need any membership. Project writes need editor or owner — a viewer gets a 403. Managing the org itself (rename/delete, members, invitations, billing) is owner-only.

  • Billing. If an org's plan has lapsed it becomes read-only and writes return 402. Inviting members additionally needs a current Team or Enterprise plan (402 otherwise).

  • Errors are readable. HTTP failures are surfaced as isError tool results with a plain message, e.g. "Forbidden (403): not a member of this organisation", "Conflict (409): a project with that name already exists", "Authentication failed (401): …".


Notes & caveats

  • update_project is a revision-safe whole-document save. The API's save endpoint is a PUT of the entire project JSON. To make partial edits safe, update_project defaults to merge=true: it fetches the current document and shallow-merges the top-level keys you supply (so {racks:[…]} replaces only the racks). Pass merge=false to replace the whole document, in which case you must provide a complete, valid project. get_project returns a revision; pass it to update_project when the edit was derived from that read. The update is refused with 412 if something else saved first. Omitting it still performs a fresh conditional read immediately before saving.

  • Shared catalogues and roles are revision-safe whole-array writes. set_org_catalog and set_org_roles replace their entire arrays. Read the matching resource first, preserve every entry you still need, and pass its returned revision to the set tool. A concurrent change is refused instead of being overwritten.

  • Live collaboration. If a project is open in a live collaboration session in the app, coordinate with the people editing it. Revision checks prevent a stale MCP save from silently overwriting a newer room save, but they cannot decide whose intended change should win.

  • Export output is truncated. A large artefact is cut off in the tool reply with an explicit marker (the byte count is always reported in full). Use the app's download for the complete file.

  • export_project never silently drops data. A placement whose deviceTypeRef matches no catalogue entry comes back under unresolved rather than vanishing; pass placeholders: true to emit it as a placeholder device type so the row still imports.

  • Schema. Documents use schemaVersion: "1" and the backend rejects unknown top-level fields. The current shape includes containers, containerTypes, deviceRoles, reviewDismissals, cabling and power data. Use the project object returned by get_project rather than the surrounding revision envelope as the update body.

For operators (publishing)

Two distribution channels, both from this mcp/ directory. Neither is done automatically — these are the manual operator steps.

npm (enables npx -y railyard-mcp and the config blocks above):

npm publish            # runs the build first via prepublishOnly; add --access public if you scope the name

package.json ships only dist/, manifest.json, README.md, INSTALL.md and LICENSE (see its files), and the prepare/prepublishOnly scripts rebuild dist/ so it is always fresh on publish. The public package name is the unscoped railyard-mcp. Publishing remains an explicit operator action; verify the version, changelog and package contents first.

Claude Desktop bundle (.mcpb, the one-click install):

npm run build                       # produce dist/
npx @anthropic-ai/mcpb pack         # bundles manifest.json + dist/ + deps into railyard-mcp.mcpb

The bundle is described by manifest.json: it declares the Node entry point and a user_config that prompts for the token (stored securely) and an optional base URL that defaults to the hosted service. Distribute the resulting .mcpb file for drag-and-drop install.

Development

npm run build      # compile once
npm run dev        # compile on change (tsc --watch)
npm test           # build and run API-contract tests
npm run test:coverage # enforce at least 80% line coverage (Node 22+)
npm run typecheck  # type-check without emitting

Source layout:

  • src/client.ts — the typed HTTP client. All auth (Authorization: Bearer), org resolution (X-Org-Id), revision preconditions (ETag / If-Match), and error mapping live here, in one place.

  • src/auth.ts — hosted/self-hosted URL validation plus the explicit cross-platform browser helper used by railyard-mcp auth.

  • src/project.ts — the canonical empty project factory used by create_project.

  • src/index.ts — the MCP server: tool definitions (zod schemas + annotations) and stdio wiring.

Available Tools

31 tools
accept_inviteAccept an invitationA

Accept an invitation addressed to the token user's email, joining that organisation at the invited role. Returns the joined org. Refused (403) if the invitation was addressed to someone else, and (402) if the organisation's plan has lapsed or been downgraded since the invitation was sent.

ParametersJSON Schema
NameRequiredDescriptionDefault
inviteIdYesThe invitation's id — from list_my_invites.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the annotations: it discloses the side effect of joining an organization, the return value ('Returns the joined org'), and specific failure modes with status codes (403 for wrong recipient, 402 for plan lapse/downgrade). This goes well beyond the readOnlyHint/openWorldHint/idempotentHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences pack the core action, outcome, return value, and error conditions with no filler. The primary purpose is front-loaded and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter action with no output schema, the description covers the key operational facts: what action is performed, who it applies to, what is returned, and what error conditions exist. No critical information needed to call the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the single required parameter inviteId is documented in the schema with a helpful source hint ('from list_my_invites'). The description does not add further parameter-level detail, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('Accept'), a specific resource ('an invitation addressed to the token user's email'), and the outcome ('joining that organisation at the invited role'). It also notes the return value, making the tool's purpose unmistakable and distinct from siblings like invite_member or revoke_invite.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use it: for invitations addressed to the authenticated user's email, via an invite id. The 403 error condition communicates a boundary (do not accept invitations meant for others), and the schema points to list_my_invites as the source of the id. It stops short of explicitly naming alternatives or saying 'when not to use', but context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

billing_manage_urlGet a Stripe billing linkA

Mint a Stripe hosted-page URL for the organisation's OWNER to open in a browser: action=subscribe opens Checkout to start a subscription, action=manage opens the Customer Portal to change the card, switch plan or cancel. This only creates a link — it does not charge anything or change the subscription; the owner completes or abandons that on Stripe's page. Requires the OWNER role and Stripe configured on the server (503 otherwise). action=subscribe conflicts (409) when a live subscription already exists — manage it instead; action=manage needs an existing billing account (400 before the first subscription).

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
actionYessubscribe = Stripe Checkout for a new subscription; manage = Customer Portal for an existing one.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond annotations: it clarifies that the tool only creates a link and does not charge or alter the subscription, and it outlines specific error conditions (503, 409, 400). Annotations only indicate readOnlyHint=false, openWorldHint=true, etc., so this is valuable extra disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a moderately long paragraph but every sentence serves a purpose: purpose, side effects, prerequisites, and error conditions. It front-loads the core action and then provides necessary caveats. The structure is effective, though slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema, the description covers essential aspects: what it returns (a URL), how to choose actions, prerequisites, and error behavior. It also benefits from annotations covering idempotency and destructiveness. Nothing critical is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds meaningful detail for the action parameter (e.g., manage = change card, switch plan, cancel), which goes beyond the schema's 'Customer Portal for an existing one.' It also implies org selection defaults, but the schema already covers that. Overall, the description enriches parameter understanding, so a 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource combination: 'Mint a Stripe hosted-page URL' and specifies the owner as the intended user. It clearly distinguishes this from sibling get_billing by focusing on URL generation rather than billing information retrieval, even without naming the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit conditions for choosing between action=subscribe and action=manage: subscribe conflicts (409) when a live subscription exists, manage requires an existing billing account (400). It also states prerequisites: OWNER role and Stripe configured (503). This is clear operational guidance, though it does not name alternative sibling tools for billing viewing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_project_nameCheck project name availabilityA
Read-only

Check whether a project name is free in an organisation (and see the URL slug it would get). Optionally exclude a project id so a rename that keeps its own name reads as available.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
nameYesThe candidate project name.
excludeNoA project id to ignore in the collision check (for renames).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the description doesn't need to restate safety. It adds behavioral value by noting the tool returns the URL slug and by clarifying that `exclude` makes a rename read as available. This goes beyond the annotations, though it doesn't cover error conditions or response format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core purpose and the most nuanced parameter behavior are both front-loaded, and every word earns its place. The description is efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema, the description indicates the tool returns the availability status and the URL slug, which is the essential information. It does not detail the exact response format (e.g., boolean vs. object), but for a simple check tool this is adequate. It also omits error conditions, but given the read-only nature and clear purpose, it is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes all three parameters with 100% coverage, so the baseline is 3. The description adds meaningful semantics for `exclude` by explaining the rename scenario, which goes beyond the schema's terse 'A project id to ignore in the collision check'. This lifts the score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Check'), a precise resource ('project name availability'), and the context ('in an organisation'), plus a notable secondary output ('URL slug it would get'). It clearly distinguishes this from sibling tools like list_projects or validate_project by focusing on availability and slug generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a concrete usage scenario for the `exclude` parameter (renames that keep their own name), which implicitly signals when to use this tool (before create or rename). It does not explicitly name alternatives or state when not to use it, but the purpose is self-evident and the exclude hint provides practical guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_orgCreate organisationA

Create a new shared organisation; the token's user becomes its owner. Returns the org's id and slug, which other tools accept as org. Note that adding members to it needs a Team or Enterprise plan (see get_billing).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName for the new organisation.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations establish this is a write operation (readOnlyHint=false) but the description adds behavioral depth: the token's user becomes owner, the response includes id and slug that other tools accept as `org`, and member addition is gated by plan. This goes well beyond what annotations and schema provide, informing the agent of side effects, output semantics, and a business-logic constraint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: core action and ownership side-effect, return value usage, and plan limitation. Information is front-loaded and there is zero repeated content from the schema or annotations. Highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one parameter and no output schema, the description covers everything needed to call the tool correctly: what it does, who gets ownership, what is returned, how the return value is used by other tools, and a plan restriction. An agent has sufficient context without additional lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the `name` parameter is already documented as 'Name for the new organisation.' The tool description adds no additional detail about parameter formatting or constraints, so the baseline of 3 applies. There is no gap requiring compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Create a new shared organisation'. It adds unique scope details (owner becomes the token's user) and is clearly distinct from sibling tools like list_orgs, rename_org, and delete_org. An agent can immediately identify this as the creation operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context that this tool is for creating a new shared organisation, with no ambiguity about when to invoke it. It also flags a relevant prerequisite by pointing to get_billing for the Team/Enterprise plan requirement, though it does not explicitly exclude alternative tools. This is slightly below 5 because it doesn't state when-not-to-use, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_projectCreate projectA

Create a new, empty project with the given name and save it to the organisation. Returns the new project's id and URL slug. Fails with a conflict if the name is already taken in the org. (Write: requires the token user to have an editor/owner role in the org.)

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
nameYesName for the new project.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate a non-read-only, non-destructive operation. The description adds valuable behavior beyond that: it requires an editor/owner role and fails with a conflict on duplicate names. These details help the agent anticipate side effects and error conditions, though it does not elaborate on the full side-effect scope (e.g., impact on existing data).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no fluff. The core action is front-loaded, followed by return value, failure condition, and a parenthetical auth note. Every sentence earns its place, and the structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a simple 2-parameter schema (1 required), full schema coverage, and annotations covering the safety profile, the description is mostly complete. It specifies the return value (id and URL slug), the conflict failure, and the auth requirement. It does not mention pagination or advanced options, but none are needed for a create operation. It could note that check_project_name exists for pre-validation, but that is optional.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters (org and name) are fully documented in the schema with defaults and meanings. The description mentions 'given name' but adds no extra semantics beyond what the schema already provides. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Create a new, empty project') and clearly differentiates from siblings like update_project, rename_project, and delete_project. It also states the scope ('save it to the organisation'), so an agent can immediately understand the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: it returns the new project's id and URL slug, fails with a conflict if the name is taken, and notes the required role. However, it does not explicitly name alternatives (e.g., check_project_name for pre-validation) or state when not to use it, so it falls short of an explicit when/when-not clause.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_orgDelete organisation (DESTRUCTIVE)A
DestructiveIdempotent

Permanently delete a shared organisation AND every project inside it. DESTRUCTIVE and NOT reversible. Requires the OWNER role. A personal organisation cannot be deleted (400). Move out any project you want to keep first (see move_project).

ParametersJSON Schema
NameRequiredDescriptionDefault
orgYesThe organisation to delete: id (org_…), slug, or name.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses that deletion is permanent and irreversible, cascades to all projects, requires OWNER role, and returns a 400 for personal organisations. This is rich behavioral context that significantly adds value over the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences front-load the destructive and irreversible nature, then cover prerequisites, an exclusion case, and a pointer to move_project. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with one parameter and no output schema, this description covers prerequisites, constraints, side effects, and the necessary preparatory action. An agent has enough information to decide correctly and avoid destructive mistakes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents the single 'org' parameter with its id/slug/name format at 100% coverage. The description does not add meaningfully new parameter semantics beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('delete') and resource ('shared organisation'), and explicitly states the cascade to 'every project inside it'. This clearly differentiates it from deleting a single project, especially with delete_project present as a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: it requires the OWNER role, cannot be used on personal organisations, and advises moving projects out first using move_project. It does not explicitly state 'use delete_project instead for a single project', so it falls just short of fully explicit alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_projectDelete project (DESTRUCTIVE)A
DestructiveIdempotent

Permanently delete a project and its entire estate. DESTRUCTIVE and NOT reversible — there is no undo. Any live collaboration sessions on it are dropped. (Write: editor/owner role required.)

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe project's id (prj_…) to delete.
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description adds crucial behavioral context: deletion is 'NOT reversible', 'there is no undo', 'live collaboration sessions on it are dropped', and a write role is required. This goes well beyond destructiveHint and readOnlyHint, and nothing in the description contradicts the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler. It front-loads the core action and scope, follows with the most important warning (permanent/no undo/dropped sessions), and ends with the authorization requirement. Every sentence contributes essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema sinyal, this description covers scope ('entire estate'), irreversibility, side effects on live sessions, and required permissions. There are only two parameters and the schema documents them fully, so the description is sufficient for an agent to invoke the tool correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents id ('prj_…') and org (org id, slug, name, defaults). The description adds the role requirement but does not need to repeat param semantics; the baseline of 3 is appropriate because the schema carries the parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Permanently delete a project and its entire estate.' This is more specific than the title and clearly distinguishes the tool from update_project, rename_project, move_project, and export_project by emphasizing permanent, irreversible deletion of the whole project scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the usage context unmistakable: use it when a project and its entire estate must be permanently removed/formatted, and it clarifies the required role ('editor/owner role required'). It does not explicitly name alternatives or state when not to use it, but the destructive-purpose framing makes the distinction from non-destructive project tools obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_projectExport projectA
Read-only

Export a project into a downstream format (see list_export_formats for the ids — e.g. nautobot-csv, netbox-csv, designbuilder-yaml, json) and return the generated files' content, plus any unresolved placements and warnings. A placement whose deviceTypeRef matches no catalogue entry is REPORTED, never silently dropped: set placeholders=true to emit it as a placeholder device type so the row still imports. Targets that need prerequisite objects return the whole ordered bundle, not just the headline table. Large files are truncated in the reply — the app's download gives the complete artefact.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
refNoA saved project's id (prj_…) or slug to export.
formatYesExport format id, e.g. "nautobot-csv". Use list_export_formats to see them.
projectNoAn inline project document to export instead of a saved one.
placeholdersNoEmit placements with no matching device type as placeholders (default false: they are skipped, but still reported).
fallbackLocationNoLocation name given to racks with no resolvable site (default "Railyard-Unassigned").

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and destructiveHint, and the description adds substantial behavior: unresolved placements are reported rather than silently dropped, placeholder emission is opt-in, prerequisite-dependent targets return the full ordered bundle, and large files are truncated in the reply. This is meaningful context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three dense sentences with no wasteful repetition. The first sentence establishes purpose and main output, the second covers a key edge case, and the third addresses truncation. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly takes responsibility for describing return values: file content, unresolved placements, warnings, bundle ordering, and truncation. Combined with full schema coverage of the six parameters, an agent has enough context to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All six parameters are documented in the schema (100% coverage), so the baseline is 3. The description adds value by clarifying that placeholders=true causes unmapped device types to be emitted as placeholders while the default still reports them, and it gives concrete format examples. This exceeds baseline without duplicating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Export a project') and clarifies it produces downstream formats while pointing to list_export_formats for the format ids. This clearly distinguishes it from get_project or list_projects, which retrieve project data rather than exporting artefacts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the relevant helper tools (list_export_formats and list_orgs) and implicitly states the use case: export a saved or inline project into an importable format. It does not explicitly state when not to use it, but the context is clear enough relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_billingGet billing stateA
Read-only

Read an organisation's plan and billing state: plan (individual/team/enterprise), status (trialing/active/past_due/canceled), seat count, trial end, current period end, whether it is currently entitled to edit (a lapsed org is read-only and its writes return 402), whether the caller may manage billing, and whether this server has Stripe self-serve configured at all. Any member may read it.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing the consequence of a lapsed org ('read-only and its writes return 402'), indicating that the tool reports whether the caller may manage billing, and noting whether Stripe self-serve is configured. This is valuable behavioral context not implied by readOnlyHint alone, and it does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient: one front-loaded sentence states the purposeament and then lists all meaningful return aspects in a compact, organized list. Every clause adds information an agent would need, with no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description effectively enumerates the content an agent can expect, including edge-case behaviors like the 402 on writes and the Stripe configuration flag. It also covers permissions and parameter defaults through context, making the tool fully approachable for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully documents the single optional 'org' parameter, including accepted forms (id, slug, or name), the default resolution order, and a pointer to list_orgs. With 100% schema description coverage, the description does not need to add parameter details, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read an organisation's plan and billing state', then enumerates the exact fields returned (plan, status, seat count, trial end, period end, entitlement, manage permission, Stripe configuration). This makes the tool's purpose unmistakable and clearly distinct from sibling tools like billing_manage_url or server_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when this tool is appropriate: it is a read-only retrieval of billing state and explicitly states 'Any member may read it'. It does not explicitly name alternative tools or exclusion conditions, but the scope is clear enough that an agent can decide to use it for reading billing rather than for generating a billing management URL.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_org_catalogGet shared device-type libraryA
Read-only

Read an organisation's shared device-type library — the catalogue of device types available to every project in the org, separate from each project's own catalogue. Returns the complete catalogue plus its revision; pass that revision to set_org_catalog. Any member may read it.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already mark this as read-only, and the description adds useful behavioral context: it returns the complete `catalogue` and its `revision`, links that revision to set_org_catalog, and states that 'Any member may read it'. This goes beyond the readOnlyHint by describing the return payload and access scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and every clause earns its place: it defines the resource, contrasts it with project catalogs, states the return value, notes the downstream tool, and gives permission information. No filler or redundancy is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description is complete: it says what the tool returns, who can use it, and how the result relates to set_org_catalog. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single optional `org` parameter is fully documented in the schema, including defaults and how to discover options via list_orgs. The description adds no additional parameter-level semantics, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read an organisation's shared device-type library'. It distinguishes itself from project-level catalogs by stating it is 'separate from each project's own `catalogue`', so an agent can tell it apart from tools like get_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly signals the tool is for the org-wide shared library rather than a project's own catalog, and it gives the follow-up step of passing the returned `revision` to set_org_catalog. It does not explicitly name an alternative tool for project-level catalogs, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_org_rolesGet shared device rolesA
Read-only

Read an organisation's shared device-role vocabulary. Returns the complete roles array and its revision; pass that revision to set_org_roles. Project-level deviceRoles are layered on top.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds value by disclosing the revision-handoff behavior and the layering relationship with project-level deviceRoles. It doesn't describe pagination or error cases, but for a simple read of a single org's roles, the description plus annotations are sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the operation and return value, the second gives the revision handoff, the third clarifies the layering with project-level roles. No filler, no repetition of schema content, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-parameter tool with no output schema, the description covers the return value, the revision handoff, and the relationship to project-level roles. The only minor gap is that it doesn't explicitly state the response shape beyond 'roles array and revision', but that is adequately specified. The tool is simple enough that nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the org parameter well, including the defaulting behavior and the pointer to list_orgs. The description adds the context that the org is the target of the role vocabulary read, which is marginal but useful. Baseline 3 is exceeded slightly because the description reinforces the parameter's role in the operation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read'), a specific resource ('an organisation's shared device-role vocabulary'), and the exact return payload ('the complete roles array and its revision'). It also distinguishes itself from the sibling set_org_roles by naming it and explaining the revision handoff. This is a clear, non-tautological purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to pass the returned revision to set_org_roles, which tells the agent when this tool is a prerequisite and how it relates to its write counterpart. It also clarifies that project-level deviceRoles are layered on top, so an agent knows this is the org-level view and not the full effective picture. This is strong routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_projectGet projectA
Read-only

Fetch a project's full JSON document and revision ETag, addressed by id or URL slug. The project contains the complete estate model, including the flexible container hierarchy, racks and placements, device roles, review dismissals, cabling and power. Pass revision to update_project when applying work derived from this response so concurrent changes are not overwritten.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
refYesThe project's id (prj_…) or its URL slug.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnlyHint=true and destructiveHint=false. The description adds that the response includes a revision ETag, lists the contents of the estate model, and notes the concurrency behavior with update_project, going beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the action and addressing, the second describes the response contents, and the third gives a workflow hint. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with annotations and no output schema, the description adequately covers what is returned, how to address the resource, and how to safely use the response with update_project. An agent has enough context to invoke and handle the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema already describes both org and ref, including the id/slug format. The description only restates that the project is addressed by id or URL slug, adding no new parameter semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Fetch') and resource ('a project's full JSON document and revision ETag'), and clarifies addressing by id or URL slug. It differentiates from siblings like list_projects and export_project by emphasizing 'full JSON document' and the revision ETag.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context by explaining that the returned revision should be passed to update_project to prevent overwriting concurrent changes. This establishes a read-modify-write use case, but it does not explicitly mention when not to use it (e.g., when a summary is sufficient, use list_projects).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

invite_memberInvite someone to an organisationA
Idempotent

Invite an email address to join an organisation at a role (default editor), and email them the invitation where the server has mail configured. Requires the OWNER role, and a current Team or Enterprise plan — a personal or Individual-plan org cannot add members (402), and neither can one whose plan has lapsed. Re-inviting a still-pending email updates its role. An address that is already a member is refused (409).

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
roleNoRole to grant on acceptance: viewer (read-only), editor (can create and change projects) or owner (full control). Defaults to editor.
emailYesThe invitee's email address.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses side effects beyond annotations: sends email when mail is configured, requires ownership and plan, re-inviting updates role, and returns 409 for existing members. This enriches the idempotentHint and readOnlyHint=false annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each earning its place: action/default, mail behavior, permission/plan constraints, and re-invite/conflict edge cases. Front-loaded with the primary purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers prerequisites, failure modes (402, 409), default role, and re-invite semantics, so an agent can call it correctly. No output schema exists, but return values are not needed for successful invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and already documents org, role, and email, so the description is not required to compensate. It adds contextual behavior such as email delivery and default role, but little new parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the exact action: invite an email address to join an organisation with a role and email the invitation. This clearly differentiates from siblings that list, revoke, or accept invitations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit preconditions: OWNER role, active Team/Enterprise plan, and notes when it will not succeed (personal/individual-plan or lapsed-plan orgs, already a member). It does not name sibling alternatives such as list_invites or revoke_invite, but gives clear context for when the action applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_export_formatsList export formatsA
Read-only

List the export targets this Railyard build supports — each format's id (what export_project takes), a one-line description, and the file extension it produces.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, and destructiveHint=false, so the agent knows this is a safe, non-destructive read operation. The description adds the behavioral detail that the result includes each format's id, description, and file extension, which is useful. However, it doesn't disclose whether the list is static or dynamic, whether it requires prior setup, or any rate-limit considerations. With annotations covering the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the main purpose ('List the export targets this Railyard build supports') and then specifies the exact output fields. Every word earns its place, and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list tool with no output schema, the description is nearly complete. It tells the agent what the tool returns (id, description, extension) and how the id relates to export_project. The only minor gap is that it doesn't explicitly state the output format (e.g., JSON array), but the absence of an output schema and the simplicity of the tool make this a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema is empty. The description correctly explains what the tool returns, which is the only semantic content needed. Since there are no parameters to document, the description doesn't need to add parameter-level detail. The baseline for 0 params is 4, and the description meets that by clearly stating the output content.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: listing export targets supported by the Railyard build. It specifies the resource (export formats), the action (list), and the key content of the result (id, description, file extension). It also distinguishes itself from the sibling export_project by noting the id is what export_project takes, which helps an agent understand its role in the workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when an agent needs to discover available export formats and their IDs before calling export_project. It doesn't explicitly state 'use this before export_project' or list alternatives, but the mention of 'what export_project takes' provides clear contextual guidance. It could be stronger with an explicit when-to-use statement, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_invitesList pending invitationsA
Read-only

List an organisation's pending invitations (email, role, when sent). Requires the OWNER role, since the list holds the addresses of people who are not members yet.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds value by disclosing that this endpoint requires the OWNER role because it exposes email addresses of people who are not yet members, which is a privacy/authorization consideration beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundant phrasing. The core function and returned fields are front-loaded in the first sentence, and the authorization requirement is explained concisely in the second sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool with one optional parameter, the description covers purpose, key output fields, and authorization requirements. It does not state the exact response structure, but there is no output schema and the tool is simple enough that this is a minor gap rather than a critical omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema describes the org parameter thoroughly, including accepted formats, default behavior, and a pointer to list_orgs. With 100% schema description coverage, the description is not required to add further parameter detail, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('an organisation's pending invitations') and specifies the data returned (email, role, when sent). It accurately distinguishes this org-level invitation list from sibling tools like list_members or list_my_invites through the phrase 'organisation's pending invitations'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states a required condition for use ('Requires the OWNER role') and explains why, which is useful. However, it does not explicitly compare against alternatives such as list_my_invites for personal invitations or list_members for current members, leaving the when-to-use vs alternatives guidance implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_membersList organisation membersA
Read-only

List an organisation's members — user id, email, name, role (viewer/editor/owner) and when they joined. Any member may read the roster. The user id is what set_member_role and remove_member take.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=true. The description adds value beyond annotations by stating that any member may read the roster (an access-control trait) and by listing the returned fields (user id, email, name, role, join date) since there is no output schema. This complements the annotation-provided safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences with zero filler. The first sentence front-loads the core purpose and output fields; the second adds access context and cross-references to sibling tools. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one optional parameter and no output schema, the description is nearly complete: it states the access rule, lists the output fields, and relates the returned user id to mutation tools. It does not mention pagination or result-size limits, which might matter for large organisations, but this is a minor omission given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema description for the 'org' parameter is thorough (org id, slug, or name, defaults, and reference to list_orgs). The tool description does not need to add parameter details. While it doesn't repeat or augment param semantics, the schema already carries the full burden, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') with a clear resource ('an organisation's members') and enumerates the exact fields returned. It also distinguishes this tool from sibling mutation tools by noting the user id is what set_member_role and remove_member take, which differentiates it from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: 'Any member may read the roster' indicates when the tool can be used (any member, read-only). It implicitly contrasts with mutation tools by naming them (set_member_role, remove_member), suggesting they are for changes. However, it does not explicitly state 'use X instead when modifying roles' or list exclusions, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_my_invitesList my pending invitationsA
Read-only

List invitations addressed to the token user's own email — organisations they have been invited to but not yet joined. Pass an id from here to accept_invite.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context about the invitation scope (own email, not-yet-joined organizations), but does not disclose return format, ordering, or behavior when there are no invites. This is acceptable given the annotations, but not richer than needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no filler. The first sentence front-loads the exact purpose and scope; the second adds a useful cross-reference to the acceptance flow. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only list tool with annotations covering safety and an output implication ('Pass an id from here'), the description is complete. An agent knows what invites are listedanding how the result feeds into accept_invite without needing further detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema and description leave nothing to clarify at the parameter level. The mention of passing an id to accept_invite concerns the output rather than an input parameter, which is appropriate for a no-parameter list tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it lists invitations addressed to the token user's own email. It also clarifies the invitations are for organizations not yet joined, and references the downstream accept_invite flow, making it clearly distinguishable from the sibling list_invites.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the use case clear: list pending invitations for the authenticated user's own email. It does not explicitly name an alternative tool such as list_invites or state when not to use this tool, but the 'own email' scoping strongly implies the distinction, and the reference to accept_invite provides actionable follow-up context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_orgsList organisationsA
Read-only

List the organisations the token's user belongs to, with each org's id, slug, role (viewer/editor/owner), plan and billing status. The id or slug is what you pass as the org argument to other tools.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the description doesn't need to state that. It adds the context that the result includes specific fields and that the id/slug is intended for use elsewhere, which goes beyond the annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The core action and returned fields are front-loaded, and the follow-up sentence adds actionable guidance about the id/slug. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with no output schema, the description covers the essential information: what it returns and how to use it. It doesn't specify the exact JSON structure (e.g., array vs. object) but that's minor and inferable. The description is adequate for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the description correctly omits parameter details. The baseline for 0 params is 4, and the description adds no unnecessary param info. It's appropriately silent on params.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List'), names the resource ('organisations the token's user belongs to'), and enumerates the returned fields (id, slug, role, plan, billing status). It clearly distinguishes this from sibling tools like list_projects or list_members by focusing on org-level data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a clear use case: obtaining org ids/slugs to pass as the `org` argument to other tools. While it doesn't explicitly mention when not to use it or name alternatives like whoami, the purpose is sufficiently scoped that an agent can infer when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsList projectsA
Read-only

List projects in an organisation (id, name, slug, last-updated time). Omit org to use the default organisation.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this read-only and non-destructive; the description adds the concrete output shape (id, name, slug, last-updated time), which is useful and not present in annotations. It does not discuss pagination or ordering, but that is a minor gap for a simple list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence front-loads the operation and immediately gives the most important usage exception. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-optional-parameter, read-only list tool with full schema coverage, this description is sufficient: it defines the scope, the output fields, and the default-org behavior. Lack of an explicit note about pagination or limits does not make it incomplete for selecting and invoking correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single org parameter is already fully documented in the schema, including its default resolution and how to discover options via list_orgs. The description restates the default-org behavior but adds no new semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('List projects') and a bounded resource ('in an organisation'), and enumerates the returned fields. The plural 'projects' clearly separates it from get_project, so an agent can select it confidently.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a practical usage condition: omit org to use the default organisation. It does not explicitly name alternatives like get_project or list_orgs, but the context is unambiguous and there are no misleading exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

move_projectMove project to another organisationA

Move a project out of its current organisation into another one the token user can write to. toOrg is the destination org (id, slug or name); org is the source org it currently lives in (defaults as usual). Requires an editor/owner role in BOTH organisations. If the name is already taken in the destination the move conflicts (409) — pass name to rename the project as part of the move; check_project_name against toOrg tells you in advance.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe project's id (prj_…) to move.
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
nameNoOptional new name, applied as part of the move (to settle a name clash in the destination).
toOrgYesDestination organisation: id (org_…), slug, or name.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag readOnlyHint=false, destructiveHint=false, openWorldHint=true, idempotentHint=false. The description adds value beyond that: it discloses the dual-role requirement (editor/owner in BOTH orgs), the 409 conflict behavior, and the idempotent-vs-conflict distinction when `name` collides. Useful behavioral context not present in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with no filler. The first sentence delivers the core action and destination constraint; the second packs the source-default semantics, permission requirement, conflict behavior, and a pointer to a sibling tool. Every clause earns its place and everything an agent needs is presented in order of importance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, permission-gated tool with 4 parameters and no output schema, the description covers itself fully: purpose, parameter roles, permissions, error behavior, and a resolution path. Nothing required to call it correctly (roles, destination resolution, name-clash handling) is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes further by clarifying the relationship between parameters: `toOrg` is the destination, `org` is the source with default behavior, and `name` is specifically for settling a destination name clash as part of the move. This adds interaction semantics beyond the standalone schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Move a project out of its current organisation into another one'), names the destination constraint ('the token user can write to'), and clearly differentiates from siblings by describing the move semantics rather than mere rename/delete. It also names the permission requirements, making the operation's scope and target unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when the move conflicts (name taken → 409), how to resolve it (pass `name`), and recommends check_project_name against `toOrg` as a pre-check. This gives clear context and one explicit alternative, though it does not formally enumerate exclusions such as 'use rename_project when staying within the same org'. Clear but not exhaustive on when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_memberRemove a member (DESTRUCTIVE)A
DestructiveIdempotent

Remove a member from an organisation, revoking their access to all of its projects and dropping their live sessions at once. Requires the OWNER role. Removing the last owner is refused (409). The subscription's seat count is reconciled afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
userIdYesThe member's user id (usr_…) — from list_members.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral detail beyond the destructiveHint/readOnlyHint annotations, such as revoking access to all projects, dropping live sessions, refusing last-owner removal, and reconciling seat counts. No contradiction with the annotations was found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each carrying necessary information: the core destructive action, the permission requirement and key failure case, and a post-condition. The most important information is front-loaded and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive member-removal tool with no output schema, the description covers the essential context: what is destroyed, prerequisites, an error condition, and side effects. The annotations and schema cover safety and parameter details, so nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents org and userId fully. The description does not need to add parameter details; it stays at the baseline without adding extra parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Remove a member from an organisation') and goes beyond the title by detailing concrete consequences: revoking project access and dropping live sessions. This clearly distinguishes it from sibling tools like set_member_role and invite_member.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: it requires the OWNER role and explicitly notes that removing the last owner is refused with a 409. It does not explicitly name alternatives for when to use set_member_role instead, but the preconditions and edge-case exclusion are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_orgRename organisationA

Change an organisation's display name. Requires the OWNER role in it.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgYesThe organisation to rename: id (org_…), slug, or name.
nameYesThe new organisation name.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the agent knows this is a mutating, non-idempotent, non-destructive operation. The description adds the permission requirement and that it only changes the display name, but it does not disclose response format, error conditions (e.g., name collision, invalid org), or side effects beyond the name change. With annotations covering the basic behavior, the description adds some value but not rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste. The primary action is front-loaded, and the permission constraint is stated immediately after. Every word carries meaning and there is no redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a simple rename operation with a fully documented 2-parameter schema and annotations covering the safety profile. However, it does not mention the return value, idempotency consequences (annotations say idempotentHint=false), or failure modes (e.g., what happens if the org doesn't exist or the caller lacks OWNER). For a mutating tool, those details would help an agent anticipate errors, but the schema and annotations already carry the core context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are individually documented in the schema, so the baseline is 3. The description adds one meaningful semantic detail: it clarifies that 'name' is the display name specifically, and it implies that the 'org' parameter accepts id/slug/name as identifiers (though this is already in the schema). The description doesn't add syntax or format details beyond the schema, but the schema itself is strong. A 4 is appropriate because the description reinforces the display-name semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+object construction ('Change an organisation's display name') and clearly distinguishes this from rename_project and other rename operations in the sibling list. It also names the exact resource (organisation) and the scope of the change (display name). An agent can select this tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states the permission prerequisite ('Requires the OWNER role in it'), which is important usage guidance. However, it does not explicitly mention when not to use this tool (e.g., use rename_project for projects, or create_org/set_org_catalog for other org changes). The sibling list makes those alternatives inferable, but the description itself stops short of naming them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_projectRename projectA

Rename a project. Updates both its display name and its URL slug together, after checking the new name is unique in the org. Returns the new slug. (Write: editor/owner role required.)

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe project's id (prj_…).
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
nameYesThe new project name.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate a write, non-destructive, non-idempotent operation. The description adds meaningful context beyond that: the display name and slug are updated atomically, uniqueness is validated within the org, editor/owner role is required, and the new slug is returned. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, followed by the key behavioral effects and the access requirement. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter mutation tool with no output schema, the description covers the action, side effects, validation, return value, and required role. The org default is already explained in the schema, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes all three parameters at 100% coverage, so baseline is 3. The description adds value by explaining that the name parameter drives both the display name and URL slug, and that uniqueness is checked in the org, which clarifies the effect of the name parameter beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action and resource ('Rename a project') and adds distinctive behavioral details: updating display name and URL slug together, checking uniqueness, and returning the new slug. This clearly differentiates it from broad tools like update_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the action ('rename'), and the role requirement is explicit. However, it does not explicitly contrast with siblings such as update_project or check_project_name, nor state when to prefer this tool over them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_inviteRevoke a pending invitationA
Idempotent

Withdraw a pending invitation so it can no longer be accepted. Requires the OWNER role. Has no effect on someone who has already joined — use remove_member for that.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
inviteIdYesThe invitation's id — from list_invites.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description adds valuable behavioral context beyond those hints, particularly the OWNER role requirement and the caveat that revoking has no effect on users who have already joined. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences with no filler. The purpose is front-loaded, followed by the role requirement and then the alternative tool for the related but different case. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter action with no output schema, the description, schema, and annotations together cover what the agent needs: purpose, role requirement, relevant edge case, parameter details, idempotency, and non-destructiveness. The only minor omission is a description of the response behavior after revocation, but that is not essential for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are fully documented in the input schema. The tool description itself adds no extra parameter-level semantics beyond what the schema already provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Withdraw a pending invitation so it can no longer be accepted.' It also differentiates the tool from remove_member, so an agent can clearly identify when this tool applies versus a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use the tool (for pending invitations) and when not to use it ('Has no effect on someone who has already joined — use remove_member for that'). It also states the OWNER role requirement, giving the agent clear authorization guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

server_infoServer infoA
Read-only

Report the Railyard server's health and capabilities: version, the project schemaVersion it understands, whether persistence (a database) and authentication are enabled, and whether Stripe billing is configured. Useful to confirm the server is reachable and which features exist.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds value by specifying exactly what information is reported (version, schemaVersion, persistence, auth, Stripe), which goes beyond the generic read-only hint. This helps an agent know what to expect from the response.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundancy. The first sentence front-loads the core purpose and content list, and the second sentence adds a practical use case. Every word earns its place, making it efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema), the description adequately covers what the tool does and what it returns. It lists the capabilities and states its utility for checking server reachability. While it doesn't describe error scenarios or detailed response structure, that's not necessary for an info tool with annotations covering safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is an empty object, so there is nothing to explain about parameters. The description correctly does not mention any, which is appropriate. With 0 parameters, the baseline is 4, and the description fully aligns with that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: reporting the Railyard server's health and capabilities, and enumerates the specific items (version, schemaVersion, persistence, authentication, Stripe billing). It distinguishes itself from sibling tools that operate on specific resources (projects, orgs, members) by focusing on server-level information, so an agent can easily tell it apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear usage context: 'Useful to confirm the server is reachable and which features exist.' This implies when to use it, but it doesn't explicitly mention when not to use it or name alternative tools. However, given the uniqueness of the tool among siblings, the context is sufficient for basic guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_member_roleChange a member's roleA
Idempotent

Change an existing member's role in an organisation. Requires the OWNER role. Demoting the last remaining owner is refused (409) — promote someone else first. The member's live collaboration sessions are dropped so they reconnect with the new role.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
roleYesviewer = read-only, editor = can create and change projects, owner = full control including members and billing.
userIdYesThe member's user id (usr_…) — from list_members.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses meaningful behavioral details: the OWNER permission requirement, the 409 refusal when demoting the last owner with a suggested workaround, and the side effect of dropping live collaboration sessions. This goes well beyond what readOnly/idempotent/destructive hints alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences front-load the primary action, then add permission requirements, a failure case, and a side effect. There is no filler or repetition of schema content, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the essential invocation context: the action, required permission, an important edge case with status code, and the consequence of the action. An agent has what it needs to call the tool correctly and anticipate the key failure mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already fully documented in the schema. The description adds context around the 'owner' role via the last-owner demotion rule, but it does not add new parameter-level meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Change an existing member's role in an organisation.' It clearly distinguishes this from siblings like remove_member (removal) and set_org_roles (org-level roles), making the tool's purpose immediately identifiable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool—when you need to change a member's role—and provides contextual constraints like requiring the OWNER role. However, it does not explicitly contrast this tool with alternatives such as remove_member or set_org_roles, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_org_catalogReplace shared device-type library (DESTRUCTIVE)A
DestructiveIdempotent

Replace an organisation's shared device-type library with the given array. DESTRUCTIVE: this is a whole-library write, not a merge — types absent from catalogue are removed. Read the current library with get_org_catalog and send it back with your additions plus its revision. The write fails rather than overwriting a concurrent change. Requires an editor/owner role.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
revisionYesThe quoted numeric ETag returned by the matching read tool, for example "7".
catalogueYesThe complete new library: a JSON array of device types, in the shape get_org_catalog returns.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

It clearly discloses the destructive semantics: types absent from catalogue are removed, this is a whole-library write, concurrent changes are protected by revision checking, and an editor/owner role is required. This goes well beyond the annotations, which only flag destructiveHint and readOnlyHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three focused sentences with the most important information front-loaded: destructive, whole-library write, not a merge. Every sentence contributes either behavioral context, workflow guidance, or role requirements.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive write tool with no output schema, the description covers the operation's effect, the required read-before-write sequence, concurrency behavior, and authorization requirements. An agent has enough context to call it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all three parameters with 100% coverage. The description adds helpful workflow context by tying revision to get_org_catalog, but it does not add substantial meaning beyond the schema definitions for catalogue, revision, and org.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Replace an organisation's shared device-type library with the given array', a specific verb and resource. It also explicitly distinguishes this as a destructive whole-library write rather than a merge, so it cannot be confused with read or update tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear workflow: read the current library with get_org_catalog, add your changes, and send back the revision. It also states the write fails rather than overwriting concurrent changes and requires an editor/owner role, but it does not explicitly enumerate when-not-to-use cases beyond noting it is not a merge.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_org_rolesReplace shared device roles (DESTRUCTIVE)A
DestructiveIdempotent

Replace an organisation's shared device-role vocabulary. This is a whole-array write: read it with get_org_roles, preserve the entries you still need, and pass back that response's revision. The write fails rather than overwriting a concurrent change. Requires an editor/owner role.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
rolesYesThe complete new shared device-role vocabulary.
revisionYesThe quoted numeric ETag returned by the matching read tool, for example "7".

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite readOnlyHint=false, destructiveHint=true, and idempotentHint=true already present in annotations, the description adds valuable context: it clarifies the destructive whole-array replacement, the revision-based optimistic concurrency (failure on concurrent change), and the authentication requirement (editor/owner role). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, front-loads the core purpose, and then covers the workflow, safety, and prerequisites without redundant filler. Each sentence adds distinct value, and the structure guides the agent from 'what' to 'how' efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has only 3 parameters, no output schema, and rich annotations, and the description covers purpose, workflow, concurrency, and auth. The only minor gap is that it does not describe what a successful write returns (e.g., the new revision), which could matter for chaining calls, but this is not critical given the agent can infer from the read tool's pattern.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter already has a description, so the baseline is 3. However, the description reinforces meaning by calling roles the 'complete new shared device-role vocabulary' and explaining that revision is the quoted ETag from the matching read tool, adding interpretive guidance beyond the schema's literal pattern.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Replace an organisation's shared device-role vocabulary') and clearly distinguishes it from the sibling read tool get_org_roles by describing it as a whole-array write. It also references get_org_roles directly, making the read/write pair unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs the agent to first read with get_org_roles, preserve needed entries, and pass back the returned revision. It also notes the concurrency protection ('fails rather than overwriting a concurrent change') and the required editor/owner role, giving clear when-to-use context and prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_projectUpdate / overwrite project (DESTRUCTIVE)A
Destructive

Save changes to an existing project via a full-document PUT. DESTRUCTIVE: the saved document REPLACES the stored one. By default (merge=true) the given project fields are shallow-merged over the current document (only the top-level keys you supply are replaced, e.g. pass just {racks:[…]} to swap the racks) — recommended. With merge=false, project is taken as the entire new document and must be a complete, valid project. The project's real id is always preserved. Concurrent browser or MCP saves are guarded by revision checks: a stale update is refused instead of overwriting newer work.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe project's id (prj_…) or slug identifying which project to update.
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
mergeNotrue (default): shallow-merge over the current doc. false: replace the whole document.
projectYesProject fields. With merge=true, a partial set of top-level keys to overwrite (e.g. {name, racks, catalogue}). With merge=false, the complete project document.
revisionNoRevision returned by get_project. When present, the update is refused if the project changed since that read. When omitted, update_project safely reads the latest revision immediately before saving.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the destructiveHint annotation by explaining exactly what is destroyed: the stored document is REPLACED, and with merge=false the supplied `project` becomes the entire new document. It also discloses that the real id is preserved and that revision checks refuse stale updates, adding valuable behavioral context about concurrency and safety.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: it front-loads the destructive warning, explains both merge modes with a concrete example, and ends with the concurrency safeguard. The structure moves from the most critical safety information to supporting parameter behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, 5-parameter tool with no output schema, the description is remarkably complete. It covers default org behavior indirectly via the schema, explains merge modes, revision safety, id preservation, and the full-document PUT semantics. An agent has everything needed to invoke the tool correctly without consulting additional documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all five parameters at 100% coverage, so the baseline is 3. The description adds meaning beyond the schema by calling shallow-merge "recommended," explaining that merge=false requires a complete valid project, and noting that the project's real id is always preserved. This extra context helps the agent choose parameter values correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: "Save changes to an existing project via a full-document PUT" and immediately clarifies the destructive semantics. It clearly distinguishes itself from sibling tools like create_project, rename_project, and delete_project by targeting existing projects and overwriting their content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool gives clear context for when it applies: updating an existing project rather than creating or deleting one. It also provides guidance on choosing merge=true vs merge=false, recommending the shallow-merge mode. It does not explicitly name alternative tools or state when not to use it, but the existing-project framing is strong enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_projectValidate projectA
Read-only

Run the current Railyard server-side soft validation over a project: rack placement/site findings and power-model problems. Structurally invalid documents, including malformed hierarchy or cabling references, are rejected before this report. Each raw problem carries a severity, stable code, message and relevant rack or placement ids. Nothing is saved; the interactive app may group or dismiss eligible review findings separately.

ParametersJSON Schema
NameRequiredDescriptionDefault
orgNoOrganisation to target: an org id (org_…), slug, or name. Defaults to RAILYARD_ORG, or to your first (personal) org if that is unset. Use list_orgs to see the options.
refNoA saved project's id (prj_…) or slug to validate.
projectNoAn inline project document to validate instead of a saved one (e.g. a draft before saving).

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even with readOnlyHint=true and destructiveHint=false already available, the description adds meaningful behavioral context: 'Nothing is saved', the precedence rule that structurally invalid documents are rejected before producing the report, the shape of each finding (severity, stable code, message, rack/placement ids), and how the interactive app may later group or dismiss findings. These go well beyond what the annotations convey and align with readOnlyHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, all information-dense and front-loaded: purpose first, then failure precedence, then output shape, then side-effect disclosure. No sentence is filler or redundant with the annotations or schema, and the whole description is proportionate to the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema, the description compensates well by describing the finding format. It covers validation scope, rejection behavior, persistence, and downstream app behavior. The only gap is that with 0 required parameters it never clarifies what happens if both 'ref' and 'project' are omitted, or that they are alternative/possibly mutually exclusive inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; each parameter (org, ref, project) is already well documented in the schema, including the inline-vs-saved distinction and defaults. The description adds context about what validation produces but no parameter-level nuance beyond the schema, so it does not push above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run the current Railyard server-side soft validation over a project') and enumerates the exact kinds of findings (rack placement/site findings, power-model problems). It clearly distinguishes this from sibling tools like get_project (read vs. validate) and check_project_name (name-only check vs. full soft validation).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the verb and resource, and the schema notes that 'project' supports drafts before saving, but the description itself gives no explicit when-to-use vs. sibling tools like check_project_name or get_project, nor any when-not-to-use guidance. It meets the 'implied usage' bar but adds no routing or exclusion context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiWho am IA
Read-only

Return the Railyard user the configured personal access token authenticates as (id, email, name). Use this to confirm the token works and which account it belongs to.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds value by explaining the token-based authentication context and the exact output fields (id, email, name), which goes beyond the annotations. It does not mention error behavior (e.g., invalid token), but for a simple read-only identity check, this is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the core purpose and then adds the usage context. Every word earns its place, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool with no output schema, the description is nearly complete. It states what the tool returns and why to use it. The only missing piece is what happens on authentication failure, but that is a minor omission given the tool's simplicity and the annotations covering safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema provides no parameter semantics to cover. The description compensates by explaining what the tool does with the configured token, which is the only relevant input. A baseline of 4 is appropriate for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the Railyard user associated with the configured personal access token, including the specific fields (id, email, name). It uses a specific verb ('Return') and resource ('the Railyard user'), and it is easily distinguished from sibling tools like list_orgs or server_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use this tool to confirm the token works and which account it belongs to. This provides clear context for when to use it, and the sibling list shows no other tool that does this, so there is no ambiguity about alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 31 tool updatesv0.2.1
    • First observedaccept_invite
    • First observedbilling_manage_url
    • First observedcheck_project_name
    • First observedcreate_org
    • First observedcreate_project
    • First observeddelete_org
    • First observeddelete_project
    • First observedexport_project
    • First observedget_billing
    • First observedget_org_catalog
    • First observedget_org_roles
    • First observedget_project
    • First observedinvite_member
    • First observedlist_export_formats
    • First observedlist_invites
    • First observedlist_members
    • First observedlist_my_invites
    • First observedlist_orgs
    • First observedlist_projects
    • First observedmove_project
    • First observedremove_member
    • First observedrename_org
    • First observedrename_project
    • First observedrevoke_invite
    • First observedserver_info
    • First observedset_member_role
    • First observedset_org_catalog
    • First observedset_org_roles
    • First observedupdate_project
    • First observedvalidate_project
    • First observedwhoami

TDQS

A4.1/5.0

Scored across 31 tools

Disambiguation5/5

Each tool maps to a distinct resource and action—org, project, member, invite, billing, catalog/roles, validation/export—so there is little risk of selecting the wrong one. The only close pair, list_invites and list_my_invites, is clearly differentiated by scope (org-owned vs token user's).

Naming Consistency4/5

The vast majority follow a consistent verb_noun snake_case pattern (list_projects, set_member_role, accept_invite). Exceptions like whoami, server_info, and especially billing_manage_url break the pattern without causing real confusion.

Tool Count2/5

At 31 tools, this sits well beyond the 25+ threshold for a single server and will require agents to manage a large option set. The broad API scope explains the size, but it is still heavier than ideal for an MCP tool surface.

Completeness5/5

The server covers full lifecycle CRUD for organisations and projects, plus membership/invitation management, org-level catalog/roles, billing, validation, and export. I don't see obvious dead ends: writes have matching reads, destructive operations have guards, and the billing flow links out to Stripe where needed.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Facilitates project management with the Linear API via the Model Context Protocol, allowing users to manage initiatives, projects, issues, and their relationships through features like creation, viewing, updating, and prioritization.
    695 npm
    6
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides access to Linear's issue tracking system through a standardized Model Context Protocol interface, allowing users to create, update, search, and manage issues, projects, and comments via natural language.
    401 npm
    1
    MIT