Sailfish Devel MCP
Provides tools for Sailfish OS development workflows, including device access over SSH, D-Bus session calls, Lipstick screenshots, touch injection, process inspection, journal logs, RPM installation, service management, user-session command execution, app launch, browser debugging, and SDK/OBS build integration.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Sailfish Devel MCPtake a screenshot from my Sailfish device"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sailfish Devel MCP
Host-side Model Context Protocol server for Sailfish OS development workflows.
This first iteration is dependency-free: it implements MCP over stdio directly with Python's standard library and exposes typed tools around commands that are easy to get wrong during day-to-day Sailfish work.
Scope
The server currently exposes tools for:
Sailfish device access over SSH
device prerequisite checks and opt-in installation of missing optional tools
defaultuser session-bus calls
Lipstick screenshots
touchscreen discovery and tap/swipe injection
combined screenshot, touchscreen discovery, and touch injection workflows
topmost window PID lookup
process map inspection
journal log reads
RPM copy/install on a device
system and user service management
short and detached user-session command execution with the configured D-Bus environment
desktop-entry application launch as the configured session user
Sailfish Browser launch/debug helpers
preflighted Docker/mb2 RPM builds through the vendored
build-sailfishoshelpercancellable local and remote Android/AppSupport build jobs
installed SDK repository metadata refresh
OBS result and build-log lookup through
oscrepository status and search under the configured git root
RPM spec metadata summaries
QML translation search and Sailfish ternary translation checks
The committed configuration template defaults are deliberately generic.
Device tools default to the placeholder SSH target root@device, the Sailfish
user-session bus at /run/user/100000/dbus/user_bus_socket, ~/git as the
local source root, ~/OBS as the OBS checkout root, and the vendored build
helper at
src/sailfish_devel_mcp/vendor/build_sailfishos.py. Put a config file at
~/.config/sailfish-devel-mcp/config.json or pass --config to provide your
real device and OBS settings. Device entries can also carry the preferred user,
architecture, and configured release label. If an installed SDK is available,
set paths.local_sdk to its sdk-chroot path; builds will use it when it has a
target matching the requested release and architecture, otherwise they fall
back to a matching tag in the third-party coderus Docker mirror. Neither a
configured device label nor mirror tag availability independently establishes
the current official SailfishOS release or SDK target.
For a device attached over USB, pass usb as the device argument to use the
configured default device, or usb:<name> when the user supplies a device
name. The MCP keeps that device's user and session metadata but connects to
192.168.2.15 without checking or storing SSH host keys. It trusts the supplied
name and does not resolve it through DNS, so no separate validation connection
is made.
Generic Python configuration code is tracked and works in a clean checkout.
Keep all machine-specific values in ~/.config/sailfish-devel-mcp/config.json
or its environment overrides. No Python bootstrap copy is needed.
Related MCP server: micro-mcp
Host prerequisites
Use a Linux host with Python 3.10 or newer. The server has no third-party
Python runtime dependencies. Detached job supervision uses POSIX process groups
and Linux /proc. The following programs are needed only for the workflows
listed; installing every optional tool is unnecessary.
Workflow | Host requirements |
Run from the checkout |
|
Install the Python package |
|
Device commands and deployment | OpenSSH |
Repository inspection |
|
QML translation search |
|
RPM builds and SDK refresh | A running Docker engine accessible to your user, |
Installed SDK builds | A configured Sailfish Platform SDK |
OBS results and logs |
|
Android/AppSupport builds |
|
Run the MCP tool sailfish_doctor to report local command availability, helper
provenance and installed SDK targets. It does not connect to devices, test Docker
daemon access, or install host packages. Install host tools through your Linux
distribution; the package names differ between distributions. Check RPM build
requirements in each project's spec file and Android requirements in its tree's
documentation.
The remote Android build host also needs sh, the selected bash or sh,
nohup, setsid, base64, date, mkdir, chmod, cat, sed, cut,
sort, tail, and sleep, plus Linux /proc. These support job bookkeeping
and cancellation; the build command may require additional tools.
Device prerequisites and optional setup
Enable Sailfish developer mode and SSH access before using device tools. SSH
authentication and root access must be configured by the device owner. Set the
device's username and session-bus address in the local configuration: older
devices can use nemo rather than defaultuser, and the user ID can differ.
Start the graphical user session before checking Lipstick or session-bus tools.
See the Sailfish developer-mode instructions
and D-Bus reference.
Feature passed to | Device requirements |
|
|
|
|
|
|
| Python >=3.6 with standard-library |
| Optional |
|
|
|
|
|
|
|
|
The prerequisite probe itself requires sh and id. dbus-send is supplied by
dbus, and gio
is supplied by glib2
in current Sailfish packaging. Base OS tools, missing applications, permissions,
and session configuration require manual repair. Command availability and a
session socket do not prove that every device API or touch input works.
Call sailfish_device_setup with install_missing omitted or false to check
a device without installing anything:
{"device": "phone", "features": ["touch", "app_launch"]}The result lists discovered executables, missing requirements, session issues,
and optional packages that could be installed. The default features are
commands, diagnostics, screenshot, touch, app_launch, install, and
services. Select input_trace or browser explicitly when needed.
To install supported missing optional tools, explicitly opt in:
{"device": "phone", "features": ["touch", "input_trace"], "install_missing": true}Setup requires a root SSH login and pkcon or zypper. It installs only missing
python3-base (for Python) and mce-tools (for evdev_trace) from the device's
existing repositories, then checks again. Package dependencies are resolved by
the device package manager. It does not add repositories, enable SSH, change
accounts or permissions, or repair base OS packages. Package availability can
vary with the configured Sailfish release. The mappings follow the
Python spec
and Sailfish touch documentation.
Installation returns a detached job by default. Check sailfish_build_status
or use sailfish-devel-jobs wait --kind build --job-id JOB_ID; the completed
status includes the final report under device_setup and its JSON result_path.
sailfish_build_cancel requests
cancellation, but interrupting SSH does not roll back a package transaction.
wait=true requests synchronous operation for controlled diagnostics.
Trust and privacy
This is local development tooling intended for a trusted MCP client running on your own workstation. It communicates over stdio and has no network listener. The client invokes tools using your host account, SSH credentials and OBS credentials. Tool annotations describe effects; they do not enforce permissions or provide an approval layer.
Device command tools can execute arbitrary programs. Their default identity is
the SSH login account, often root; setting the session environment does not
change that identity. Select run_as_user=true when execution should use the
configured Sailfish user. Android builds intentionally accept shell commands.
Configured aliases are convenient names, not a host allowlist: explicit
user@host targets are also supported. Target strings cannot contain SSH
options or whitespace.
RPM builds and SDK maintenance use privileged Docker containers. Installed SDK
builds mount the host user's home directory and the SDK read-write. Build only
trusted source and images, including the optional third-party SDK mirror.
Repository path checks constrain direct file operations; they do not sandbox
build scripts or arbitrary device commands. The explicit chmod permission
fallback broadens file permissions; prefer ACL support and the default error
behaviour when ACL tools are unavailable.
Non-USB SSH connections follow the configured SSH host-key policy. USB mode
connects to 192.168.2.15 and deliberately disables host-key checking and storage
to support changing devices at that address. Use it only for a device connection
you trust.
By default, SSH and SCP use ~/.ssh/config when it exists, or -F /dev/null
when it is absent. Both choices skip the system SSH configuration; OpenSSH's
built-in defaults, including normal non-USB host-key checks, still apply.
An explicit paths.ssh_config must name an existing file; a missing explicit
file remains an error. This policy applies to device and Android connections.
Tool output is returned to the MCP client and may enter a model's context. Effective configuration includes private hostnames, usernames, source paths and OBS mappings. Screenshots, journals, source matches and job logs can contain private data; commands can contain credentials supplied by callers. Keep real configuration, credentials, task notes, screenshots and runtime logs out of source releases and public bug reports. The server's request logs record argument names rather than tool argument values, but job logs and verbose results retain command/output details. Configure the client's tool approvals to suit your workflow.
Running
From a checkout:
/path/to/sailfish-devel-mcp/bin/sailfish-devel-mcpTo inspect the effective configuration:
/path/to/sailfish-devel-mcp/bin/sailfish-devel-mcp --dump-configExample MCP client configuration:
{
"mcpServers": {
"sailfish-devel": {
"command": "/path/to/sailfish-devel-mcp/bin/sailfish-devel-mcp"
}
}
}The wrapper writes startup, shutdown, MCP request, and tool-call logs to
~/.local/state/sailfish-devel-mcp/server.log by default. Override the log
directory with SAILFISH_DEVEL_MCP_LOG_DIR. If the MCP client uses a different
server label, set SAILFISH_DEVEL_MCP_SERVER_LABEL in the client environment so
the wrapper startup line includes the same label.
Configuration
Example:
{
"default_device": "phone",
"devices": {
"phone": {
"ssh_target": "root@phone",
"username": "defaultuser",
"architecture": "aarch64",
"release": "live",
"user_bus_runtime_dir": "/run/user/100000",
"user_bus_address": "unix:path=/run/user/100000/dbus/user_bus_socket"
}
},
"paths": {
"git_root": "~/git",
"obs_root": "~/OBS",
"ssh_config": "~/.ssh/config",
"local_sdk": "/srv/mer/sdks/sfossdk/sdk-chroot",
"osc_api_alias": "primary",
"obs_servers": {
"primary": "primary",
"community": "community"
}
},
"default_android_build_host": "android-builder",
"android_build_hosts": {
"android-builder": {
"ssh_target": "user@android-build-host"
}
}
}Tool Notes
The server keeps local paths scoped to the configured git root, OBS checkout
root, and /tmp for tools that read or write files. Device access still uses
SSH, so the usual SSH prompts, permissions, and command failures are surfaced
as tool results.
Mutating tools are annotated as non-read-only:
sailfish_device_setup(conservatively annotated because installation is optional)sailfish_device_lipstick_screenshotsailfish_device_touchsailfish_device_touch_workflowsailfish_device_user_bus_callsailfish_device_user_session_commandsailfish_device_command_startsailfish_device_command_cancelsailfish_device_install_rpmsailfish_device_restart_servicesailfish_device_app_launchsailfish_device_browser_launchsailfish_build_rpmsailfish_build_cancelsailfish_android_buildsailfish_android_build_cancelsailfish_sdk_refresh_metadata
sailfish_device_lipstick_screenshot defaults to
~/Pictures/Screenshots/lipstick-<timestamp>.png under /home/<username>,
where username comes from the device config. Lipstick rejects screenshot save
paths outside the home directory. When that directory is missing, the tool
derives ownership with stat -L, creates Pictures as that owner/group with
mode 775, and creates Pictures/Screenshots as that owner and the
privileged group with mode 755 when the group exists.
sailfish_device_touch supports discover, tap, and swipe. Discovery
prints /proc/bus/input/devices and can also run evdev_trace -i when
include_evdev_trace is true. Tap and swipe auto-select a likely touchscreen
input event device unless input_device is supplied. Coordinates are raw
input/display coordinates, so pair this tool with a current screenshot when
choosing points.
sailfish_device_touch_workflow wraps the common UI-debugging sequence:
capture a Lipstick screenshot, optionally list touch devices, inject a tap or
swipe, and optionally capture a second screenshot. It returns each step's
structured result separately.
sailfish_device_user_session_command runs an argv command with
XDG_RUNTIME_DIR and DBUS_SESSION_BUS_ADDRESS set from the configured device.
Set run_as_user when the command should execute as the configured Sailfish
username through runuser or su. Command tools return compact metadata and
bounded output by default; set verbose only when the full command, stdout, and
stderr are needed. This synchronous tool is limited to 15 seconds.
Use sailfish_device_command_start for GDB, tracing, or any command that might
run longer than 15 seconds. It starts a detached host-side SSH process and
returns a job id immediately, so the command survives the MCP request and stdio
server exiting. Use the CLI waiter below for long waits. For occasional checks, use
sailfish_device_command_status with lines=0 and at most 15 wait_seconds. Terminal
failures include a bounded log excerpt automatically. Full logs remain under
the local MCP state directory in device-commands/<job-id>/command.log. Use
sailfish_device_command_cancel to ask the owning supervisor to terminate the
SSH process group.
sailfish_device_install_rpm accepts either rpm_path for one local RPM or
rpm_paths for a dependency set. With rpm_paths, the tool copies every RPM to
the remote remote_dir and runs one pkcon install-local or rpm -Uvh
command with all copied files, so dependencies can be resolved together.
sailfish_device_app_launch launches an installed desktop entry using gio,
as the configured device user from their passwd home directory. For example:
{"device": "my-device", "desktop_file": "sailfish-browser.desktop"}A basename resolves under /usr/share/applications; an absolute device path
ending in .desktop is also accepted. No URL is required. It does not kill
running apps, restart boosters, or change preferences. The process detaches
with stdin/stdout/stderr redirected, avoiding an SSH timeout when app children
retain the connection. wait_seconds (0–5, default 1) checks for immediate
launcher failure; timeout (up to 15 seconds) must exceed it.
The receipt includes submitted, the launch supervisor PID (not the app PID),
launch_exit_code if already available, and private remote_log/remote_status
paths under the user's ~/.cache/sailfish-devel-mcp/launch.* directory. Read those
paths with device commands for later failures and remove that specific directory
when its evidence is no longer needed. Submission or gio exit zero does not
prove foreground activation/rendering: check Lipstick's topmost PID and a
screenshot separately. This desktop-entry path was validated manually for Browser
on a Sailfish device; it does not claim to execute all of Lipstick's icon/switcher logic.
sailfish_device_browser_launch stops the browser booster service and stale
browser/firejail PIDs when requested, launches Sailfish Browser through
invoker with the display and user-session environment, then queries Lipstick
for the topmost PID and checks whether that process has libxul.so mapped.
sailfish_build_rpm can use a configured device's architecture and release
as defaults when the call includes device. If paths.local_sdk is set, the
build defaults to live and first checks the installed SDK targets. live uses
the unversioned local target for the requested architecture, for example
aarch64. A named production release is used only when passed explicitly
through the tool arguments, environment, or device config; it uses a matching
versioned local target when available, and otherwise falls back to a matching
tag in the third-party coderus/sailfishos-platform-sdk Docker mirror. Tags in
that mirror indicate image availability; they do not identify the current
official SailfishOS release or SDK target. The wrapper image defaults to
sailfish-sdk-build-engine:$USER and can be overridden with
SAILFISH_SDK_BUILD_ENGINE_IMAGE.
Use sailfish_build_preflight to validate backend, image/target selection,
architectures, local RPM inputs, pull policy, VCS behavior, and artifact paths
without pulling an image or changing the project. sailfish_build_rpm accepts
the same backend, local_sdk, target, pull_policy, no_vcs_apply, and
allow_untrusted_rpms controls. An explicitly selected local backend does
not silently fall back to Docker. Asynchronous jobs use confined UUID job
directories and expose helper metadata and RPM paths through
sailfish_build_status; use sailfish_build_cancel to terminate the tracked
process group. Successful builds run the helper in quiet mode while retaining
the complete job log. Let sailfish_build_status wait mechanically with
wait_seconds and leave lines at 0 during routine monitoring. A terminal
failure automatically includes a bounded diagnostic excerpt; request log lines
or verbose output only when more detail is needed. Status waits are capped at
15 seconds so a monitoring request stays comfortably below the stdio
transport's observed lifetime limit.
Local SDK builds use the canonical helper. Ordinary builds share the base's
mb2-managed .default working snapshot; custom repositories, packages and local
RPMs use isolated originals. snapshot_key explicitly isolates a project even
without custom inputs. Supply it when your project or machine policy requires
per-project isolation. snapshot_repository and snapshot_package use the same
helper options; repeat these inputs on later builds so outdated snapshots can
be restored. Only original targets go to mb2; working .default children are
managed by mb2. Bases remain free of project dependencies.
sailfish_android_build starts a remote Android/AppSupport build on the
configured build host. Configure android_build_hosts.<host>.project_dir or
pass project_dir to point at the remote Android source tree. The tool writes
job state under the remote state_dir, atomically creates a per-job directory,
and starts its own process group with nohup and setsid, so the SSH session
used to launch the job can disconnect without killing the build. Set
build_timeout for a remote build lifetime limit. Poll with
sailfish_android_build_status; omit job_id to list recent jobs, or pass a
job id and wait_seconds to wait mechanically for a state change. Running jobs
omit their log by default, while failed jobs include a bounded diagnostic
excerpt. Set lines or verbose only for explicit log or command inspection.
Individual status waits are capped at 15 seconds. Use
sailfish_android_build_cancel to terminate the identity-checked process group.
sailfish_sdk_refresh_metadata starts a detached job to refresh zypper metadata in the installed SDK
main target, for example aarch64.default, using the same privileged Docker
wrapper style as local SDK builds. Use it when local SDK builds fail because a
package listed in repository metadata cannot be downloaded.
sailfish_obs_results and sailfish_obs_buildlog accept a friendly server
name configured in paths.obs_servers; each value is an .oscrc alias or API
URL passed to osc -A. Omit server to use paths.osc_api_alias. The advanced
api_alias argument accepts any raw osc -A alias or API URL and cannot be
combined with server. Keep machine- or organization-specific server mappings
in the local config rather than in the repository.
sailfish_obs_results can wait mechanically in intervals of at most 15 seconds
until OBS no longer reports an active state. sailfish_obs_buildlog returns a
bounded tail by default; set verbose only when the complete response is
required.
sailfish_obs_buildlog defaults to osc api with nostream=1 so a build log
request does not become a long-running live stream. Set nostream to false
to use osc remotebuildlog.
Read-only tools include build preflight/status, the journal, topmost PID, process maps, OBS lookup, repo search, spec summary, and QML checks.
Smoke Test
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"smoke","version":"0"}}}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}' \
| /path/to/sailfish-devel-mcp/bin/sailfish-devel-mcpDevelopment
Run tests without installing the package:
PYTHONPATH=src python3 -m unittest discover -s testsThe canonical helper lives in the build-sailfishos skill. Update the exact
vendored copy with:
python3 scripts/update_build_helper.py /path/to/build-sailfishos/scripts/build_sailfishos.pyCompact jobs and diagnostics
RPM installation and SDK metadata refresh now return detached build-job receipts
by default. Wait for sailfish_build_status to finish successfully before dependent
actions; sailfish_build_cancel cancels these jobs. Explicit wait=true retains
legacy synchronous behavior. For a multi-RPM transaction, timeout bounds the whole
async job; cancellation of local SSH does not guarantee rollback of a remote RPM
transaction. Inspect package state before retrying an interrupted install.
sailfish_obs_watch runs osc results --watch --fail-on-error in a detached job,
using the selected OBS alias and package filter. It does not modify OBS sources.
For long jobs use one CLI process, without repeated model-driven status calls:
bin/sailfish-devel-jobs wait --kind build --job-id JOB_ID
bin/sailfish-devel-jobs wait --kind device --job-id JOB_ID
bin/sailfish-devel-jobs wait --kind android --host HOST --job-id JOB_ID
bin/sailfish-devel-jobs status --kind build --job-id JOB_ID
bin/sailfish-devel-jobs log --kind build --job-id JOB_ID --offset 0 --max-bytes 6000The waiter emits one JSON result at completion or failure. Exit codes are 0 for
success, 1 for failure/cancellation and 124 for wait expiration (which leaves the
job running). Installed packages expose sailfish-devel-jobs on PATH. Android
jobs accept an optional --state-dir matching their creation call.
Status results include a revision. Send after_revision on an occasional later
check to receive a minimal unchanged receipt when nothing changed. Explicit
lines, verbose, or log_offset requests still return their requested detail.
Build/device log pages are capped at 12,000 bytes and return next_offset and
eof. Tails are byte bounded too, including logs containing very long lines.
Verbose status retains full job metadata; compact completion exposes helper
version, backend, targets, duration, RPM paths, rpmlint and dependency diagnostics.
sailfish_doctor checks configured helper provenance, local executable availability,
SDK targets and configuration presence without contacting devices or OBS. It does
not prove Docker daemon access or network connectivity. The canonical helper also
provides --doctor and --refresh-metadata --target TARGET [--force-refresh].
The QML checker lexes strings/comments and multiline ternaries. It reports missing
branch-local //: and //% comments; it is a focused checker, not a full QML parser.
Maintaining the helper and skills
The vendored helper is pinned by vendor/build-helper.json with its version,
SHA-256 and available source revision/dirty state. Verify parity without editing:
python3 scripts/update_build_helper.py /path/to/build-sailfishos/scripts/build_sailfishos.py --checkThe version check, CLI-control tests and snapshot tests guard the interface. Update
the canonical skill first, vendor it, and run both test suites before release.
skills/sailfish-devel is a separate optional skill for deployment, debugging,
Android and OBS workflows. Install/symlink that directory in your skill directory;
the canonical RPM skill stays focused on building and packaging.
Available Tools
35 toolssailfish_android_buildStart Android BuildA
Start a remote Android/AppSupport build under nohup on the configured build host; project_dir must be configured for the host or supplied.
| Name | Required | Description | Default |
|---|---|---|---|
| host | No | Configured build host alias or ssh target. | |
| shell | No | bash | |
| job_id | No | Optional stable job id; autogenerated when omitted. | |
| timeout | No | Command timeout in seconds. | |
| state_dir | No | Remote directory where job state and logs are stored. | |
| project_dir | No | Remote Android tree. Defaults to the host config. | |
| build_timeout | No | Remote build lifetime in seconds; zero disables the build timeout. | |
| shell_command | Yes | Build command run from project_dir, for example 'source build/envsetup.sh && lunch ... && m ...'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, non-idempotent, open-world, non-destructive behavior, so the bar is lower. The description adds real value beyond them by disclosing that the build runs detached under nohup on a remote host, implying a fire-and-forget start that must be tracked separately, plus the project_dir configuration dependency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and resource before the host/mode and prerequisite clauses. It is compact with no filler, though the prerequisite clause is packed in and slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description could usefully explain that the call returns a job id (job_id param is described as stable/autogenerated) and how to follow up via sailfish_android_build_status. It covers the start semantics and project_dir dependency but omits the async lifecycle linkage an agent needs to complete the task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (88%), so the schema already documents nearly all eight parameters, including the shell_command example and the project_dir default. The description only reinforces the project_dir prerequisite, adding no new syntax or format detail, so the baseline 3 stands.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Start) and resource (remote Android/AppSupport build) and adds the execution mode (under nohup on a configured build host). It is clearly distinguishable from the sibling status/cancel tools by the 'Start' verb, but it never names those siblings to sharpen the routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The prerequisite 'project_dir must be configured for the host or supplied' gives conditional context for a successful call, which is useful. However, it offers no explicit when-to-use versus alternatives (e.g. when to reach for sailfish_build_rpm or the preflight tool) and no when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_android_build_cancelCancel Android BuildC
Cancel a remote Android/AppSupport build after verifying its process identity.
| Name | Required | Description | Default |
|---|---|---|---|
| host | No | Configured build host alias or ssh target. | |
| job_id | Yes | ||
| timeout | No | Command timeout in seconds. | |
| state_dir | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readOnly, non-destructive, non-idempotent, openWorld. The description adds one meaningful behavioral fact — identity verification before cancelling — which explains why a stale job_id may fail. It omits auth/SSH requirements, what happens to partially built artifacts, and the effect of the timeout, so it only modestly exceeds the annotation floor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the verb and scope front-loaded and no filler. It is efficient, though arguably too terse given the gaps it leaves unfilled.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutation tool with no output schema, 50% parameter coverage and a required job_id that is undocumented in both schema and description. The agent lacks enough to reliably construct a call or predict failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: host and timeout are documented, but job_id (the sole required parameter) and state_dir have no descriptions anywhere. The description does not compensate by explaining job_id format, how it is obtained, or the purpose of state_dir.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Cancel) plus qualified resource (remote Android/AppSupport build), which is more precise than the title alone. It does not distinguish itself from the sibling sailfish_build_cancel, so an agent must infer the Android/AppSupport scope is the differentiator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance. With two nearly identical cancel siblings (sailfish_build_cancel, sailfish_android_build_cancel) and a status sibling, the description should tell the agent which cancel path applies and whether a build must be in progress.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_android_build_hostsAndroid Build HostsBRead-onlyIdempotent
List configured remote Android/AppSupport build hosts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds no further behavioral context such as return format, pagination, authentication needs, or rate limits. It merely restates that the tool lists hosts, providing no transparency beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no waste. It efficiently conveys the tool's purpose without unnecessary detail, which is appropriate for a simple no-parameter list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema, safety annotations present), the description provides enough context to understand its purpose. It could be slightly improved by clarifying what a 'build host' represents or what the output includes, but it is complete for the basic operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema carries no parameter documentation burden. Per the rubric, zero parameters sets a baseline of 4, and there is no additional parameter semantics for the description to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('configured remote Android/AppSupport build hosts'). It is clear what the tool does, but it does not explicitly differentiate itself from siblings like sailfish_android_build or sailfish_android_build_status; an agent must infer that this tool lists hosts rather than performing builds or checking build status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives. It does not mention prerequisites, typical scenarios, or any alternatives such as sailfish_android_build or sailfish_android_build_status. The agent is left to infer usage from the tool name and context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_android_build_statusAndroid Build StatusARead-onlyIdempotent
List remote Android/AppSupport build jobs or read one job's status and recent log output.
| Name | Required | Description | Default |
|---|---|---|---|
| host | No | Configured build host alias or ssh target. | |
| lines | No | Trailing log lines to include. The default returns no running log; failed jobs include a bounded failure excerpt. | |
| job_id | No | Job id returned by sailfish_android_build. Omit to list jobs. | |
| timeout | No | Command timeout in seconds. | |
| verbose | No | Include full command, stdout, and stderr details for diagnostics. | |
| state_dir | No | Remote directory where job state and logs are stored. | |
| wait_seconds | No | Mechanically wait at most 15 seconds for a terminal state. | |
| poll_interval | No | Seconds between status checks while waiting. | |
| after_revision | No | Last revision; unchanged results omit repeated details. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds that output includes recent log output and that it reaches a remote host, which is modest extra context, but it says nothing about the polling/wait behavior or failure semantics beyond what the schema already documents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler that conveys both the list and single-read modes. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nine-parameter tool with no output schema, the description is thin: it does not convey the wait/poll semantics, the meaning of after_revision's de-duplication, or the shape of a status result. The rich schema compensates for parameter gaps, but return-value behavior remains under-described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters (including lines, wait_seconds, poll_interval, after_revision) are already documented in the schema. The description adds no syntax, format, or defaulting detail beyond that baseline, so 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: listing remote Android/AppSupport build jobs, or reading one job's status plus recent log output. The dual-mode framing distinguishes it from the mutation siblings sailfish_android_build and sailfish_android_build_cancel, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The two operating modes (list vs. read one job) are implied by the phrasing, and the schema's job_id note says to omit it to list. However, the description never states when to reach for this versus sailfish_build_status or how it fits into a build-then-poll workflow, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_build_cancelCancel Sailfish BuildB
Request cancellation of a local asynchronous Sailfish RPM build job.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description's word 'Request' usefully signals that cancellation is asynchronous and may not be immediate, but it omits auth requirements, behavior for already-finished jobs, and how to confirm cancellation. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler, front-loading the action and the scoped resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema and adequate safety annotations, the description is minimally viable. It still leaves the agent guessing where job_id comes from and what a successful 'request' implies, so it is adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single job_id parameter has no description; the description does not explain what job_id is or that it comes from sailfish_build_rpm. With one opaque required parameter, the description needed to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (cancel) and a precisely scoped resource (local asynchronous Sailfish RPM build job), which distinguishes it from sailfish_android_build_cancel and sailfish_device_command_cancel. It does not name those siblings directly, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus sailfish_build_status, sailfish_build_rpm, or the android/device cancel siblings, nor any prerequisites (e.g., the job must still be running, where job_id originates). The only hint is the implicit 'local' scoping.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_build_preflightPreflight Sailfish BuildBRead-onlyIdempotent
Validate and return a structured Sailfish RPM build plan without mutating the project.
| Name | Required | Description | Default |
|---|---|---|---|
| arch | No | ||
| clean | No | ||
| debug | No | ||
| device | No | Optional configured device to supply default release and architecture. | |
| target | No | Registered non-snapshot local SDK base target; the build runs in a managed project snapshot. | |
| backend | No | auto | |
| no_pull | No | ||
| release | No | ||
| timeout | No | Command timeout in seconds. | |
| local_sdk | No | Installed sdk-chroot path under /srv/mer; defaults to paths.local_sdk. | |
| all_arches | No | ||
| pull_policy | No | always | |
| no_vcs_apply | No | Omit for backend default; false explicitly enables VCS application. | |
| project_path | Yes | ||
| snapshot_key | No | Stable logical SDK snapshot key (for example one shared by all worktrees of an ESR generation). Omit to share ordinary builds; custom inputs infer an ESR-aware key. Supply a project key when repository policy requires isolation. | |
| artifacts_dir | No | ||
| local_rpms_dir | No | ||
| snapshot_package | No | Additional package to prepare in the project snapshot; may be repeated in the request. | |
| permission_fallback | No | error | |
| snapshot_repository | No | Additional [ALIAS=]URL repository to prepare in the project snapshot; may contain multiple values. | |
| allow_untrusted_rpms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description reinforces 'without mutating' and adds that a structured plan is returned, but it does not describe permissions, network access, or other runtime behavior beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. It clearly states the action, result, and key constraint in one pass.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 21 parameters, low schema coverage, and no output schema, the description is too sparse to fully orient an agent. It says a structured plan is returned but does not cover parameter semantics or return shape, leaving substantial gaps for a complex tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 21 parameters with only 38% description coverage, and the description mentions no parameter at all. For a tool with this many configuration options, the description provides no guidance on what any parameter means or how to use them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: validate and return a structured Sailfish RPM build plan. The phrase 'without mutating the project' distinguishes it from sibling tools like sailfish_build_rpm, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The term 'Preflight' implies use before an actual build, and 'without mutating' suggests a safe inspection step. However, the description does not explicitly state when to use this tool instead of sailfish_build_rpm or other alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_build_rpmBuild Sailfish RPMB
Start an asynchronous RPM build. Ordinary local builds share mb2 snapshots; custom inputs or an explicit key use an isolated original target. paths.local_sdk defaults to the live installed SDK, with third-party coderus Docker-image fallback only for explicit named releases.
| Name | Required | Description | Default |
|---|---|---|---|
| arch | No | ||
| wait | No | Wait for the build and return the old synchronous command result. | |
| clean | No | ||
| debug | No | ||
| device | No | Optional configured device to supply default release and architecture. | |
| target | No | Registered non-snapshot local SDK base target; the build runs in a managed project snapshot. | |
| backend | No | auto | |
| no_pull | No | ||
| release | No | ||
| timeout | No | Command timeout in seconds. | |
| local_sdk | No | Installed sdk-chroot path under /srv/mer; defaults to paths.local_sdk. | |
| all_arches | No | ||
| pull_policy | No | always | |
| no_vcs_apply | No | Omit for backend default; false explicitly enables VCS application. | |
| project_path | Yes | ||
| snapshot_key | No | Stable logical SDK snapshot key (for example one shared by all worktrees of an ESR generation). Omit to share ordinary builds; custom inputs infer an ESR-aware key. Supply a project key when repository policy requires isolation. | |
| artifacts_dir | No | ||
| local_rpms_dir | No | ||
| snapshot_package | No | Additional package to prepare in the project snapshot; may be repeated in the request. | |
| permission_fallback | No | error | |
| snapshot_repository | No | Additional [ALIAS=]URL repository to prepare in the project snapshot; may contain multiple values. | |
| allow_untrusted_rpms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds useful behavioral context beyond annotations: it is asynchronous, ordinary builds share mb2 snapshots, custom inputs or explicit keys use an isolated original target, and local_sdk defaults to the live installed SDK with a coderus Docker-image fallback only for explicit named releases. These details help an agent understand execution nuances.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and then adds two dense but relevant sentences. It is appropriately sized and contains no obvious filler, though the snapshot and SDK-fallback details are packed with jargon that could be clarified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 22-parameter build tool with no output schema, the description is too sparse. It does not explain the required project_path, how to monitor the asynchronous build (e.g., via sailfish_build_status), expected return values, or most configuration options. It leaves major gaps for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 41%, so the description must compensate but does not. It mentions paths.local_sdk defaults and hints at snapshot_key behavior via 'explicit key,' but it leaves most of the 22 parameters undocumented, including arch, clean, debug, backend, release, all_arches, artifacts_dir, and many others.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Start an asynchronous RPM build.' This clearly identifies the tool's function as initiating an RPM build asynchronously. It does not explicitly differentiate from sibling tools like sailfish_build_preflight, sailfish_build_status, or sailfish_build_cancel, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives some behavioral context about snapshot sharing and isolated targets for custom inputs, but it never states when to use this tool versus alternatives such as sailfish_build_preflight or sailfish_build_status. There are no explicit when-to-use or when-not-to-use guidelines, and no alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_build_statusBuild Job StatusARead-onlyIdempotent
Read status and recent log output for asynchronous Sailfish RPM build jobs.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | No | Trailing log lines to include. The default returns no running log; failed jobs include a bounded failure excerpt. | |
| job_id | No | Build job id returned by sailfish_build_rpm. Omit to list recent jobs. | |
| verbose | No | Include full command, stdout, and stderr details for diagnostics. | |
| max_bytes | No | ||
| log_offset | No | Read a bounded log page from this byte offset. | |
| wait_seconds | No | Wait at most 15 seconds for job completion. | |
| after_revision | No | Last revision; unchanged results omit repeated details. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds only that jobs are asynchronous and that a log excerpt is returned — useful framing, but it doesn't disclose permissions, rate limits, or polling semantics beyond what the schema already states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with zero padding, front-loading the verb and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description names the two things returned (status and recent log output), which is enough orientation for a tool whose rich parameter schema already documents paging, byte caps, and wait behavior. Slightly more on return shape (e.g., status values) would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents lines, job_id, verbose, log_offset, wait_seconds, and after_revision (max_bytes being the only bare one). The description adds no parameter-level meaning beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('status and recent log output for asynchronous Sailfish RPM build jobs'). The 'Sailfish RPM' qualifier plus 'asynchronous' distinguishes it from sailfish_android_build_status, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'asynchronous' implies this is the polling endpoint after a build is launched, and the schema notes job_id comes from sailfish_build_rpm, but the description itself gives no explicit when/when-not or routing between status, cancel, and obs tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_app_launchLaunch a Sailfish applicationA
Launch an installed .desktop entry using gio as the configured session user from their home. Detaches all streams and returns a launch receipt and remote log/status paths. Does not stop apps or boosters. Submission is not proof of foreground activation or rendering; verify separately.
| Name | Required | Description | Default |
|---|---|---|---|
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | ||
| desktop_file | Yes | Installed desktop basename (for example sailfish-browser.desktop) or absolute device path ending in .desktop. | |
| wait_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the generic write/open-world/idempotent profile; the description adds genuinely new behavioral facts: streams are detached (fire-and-forget), a launch receipt plus remote log/status paths come back, and the tool will not stop running apps or boosters. The activation caveat also warns the agent about a failure mode the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, no filler, with the primary action statement first and the caveats last. Every clause carries information (mechanism, return shape, exclusions, verification).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully names the return artifacts (launch receipt, remote log/status paths) and the non-foreground-activation caveat. The remaining gap is that the timing parameters (timeout, wait_seconds) that govern launch behavior are left entirely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: desktop_file and device are described, but timeout and wait_seconds are documented nowhere. The description does not compensate for those two parameters or explain how wait_seconds interacts with timeout, which is exactly the kind of timing semantics an agent needs here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (launch), a specific resource (installed .desktop entry), and the exact mechanism (gio as the configured session user from their home). This clearly separates it from sailfish_device_browser_launch and the command/journal siblings without needing to read any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit exclusion ('Does not stop apps or boosters') and a verification caveat ('Submission is not proof of foreground activation or rendering; verify separately'), which tells the agent when this tool's result is insufficient. It does not, however, name a concrete alternative tool when the agent actually needs to stop or confirm an app.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_browser_launchLaunch Sailfish BrowserB
Stop stale browser state, launch Sailfish Browser with display/session env, and report topmost PID details.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | Command timeout in seconds. | |
| stop_stale | No | ||
| wait_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false. Description adds context: it stops stale state (potentially destructive side effect) and reports PID details. It hints at display/session env requirements. However, it doesn't clarify auth needs, rate limits, or what 'stale state' means concretely. With annotations carrying the safety profile, this adds moderate useful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single efficient sentence, front-loaded with the main actions. No wasted words, though it crams three distinct operations into one line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, so description needn't cover return values, but it doesn't mention the PID report format. For a 5-param tool with 40% schema coverage and mutation behavior, more detail on side effects and parameters would help. Adequate but gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 40%, so description must compensate. It mentions 'display/session env' and 'topmost PID details' but does not explain url, device, timeout, stop_stale, or wait_seconds semantics. The schema partially documents device and timeout; stop_stale and wait_seconds are undocumented. Marginal value added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific actions: stopping stale browser state, launching Sailfish Browser with display/session env, reporting topmost PID. This distinguishes it from generic siblings like sailfish_device_app_launch or sailfish_device_topmost_pid, though sibling differentiation is only implied, not named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this vs. sailfish_device_app_launch, sailfish_device_command_start, or sailfish_device_topmost_pid. Usage is only implied by the action description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_command_cancelCancel Device Command JobB
Request cancellation of a detached device command and its process group.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The word 'Request' usefully signals this is a best-effort/asynchronous cancellation rather than a guaranteed kill, adding meaning beyond the readOnlyHint=false / destructiveHint=false annotations. It also discloses the scope (whole process group). However, it says nothing about permissions, reversal, or that non-idempotent behavior means repeat calls may fail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and no filler. It is appropriately sized, though it leaves room for the missing usage and parameter context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description need not explain return values, but for a cancellation tool it omits how to obtain job_id, whether the request is asynchronous, and how to confirm success. Annotations cover the safety profile, so the remaining gaps are moderate rather than severe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions job_id, where it comes from (presumably command_start), or its format. The phrase 'a detached device command' loosely implies the identifier refers to that job, but the single required parameter is effectively undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (cancel) and resource (detached device command and its process group), which separates it from siblings like sailfish_device_command_start and _status. It does not explicitly name those siblings, but the resource is distinctive enough to identify the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus sailfish_device_command_status or the build/android cancel variants, nor any precondition (e.g., that the job must be in a running state). Usage is only implied by the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_command_startStart Device Command JobA
Start a detached SSH command with the configured Sailfish user-session environment and immediately return a job id. Use this for GDB, tracing, and any device command that may exceed 15 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| command | Yes | Command argv to run; shell parsing is not applied. | |
| timeout | No | Detached job lifetime in seconds. | |
| run_as_user | No | Run through runuser/su as the configured device username. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false (mutation) and idempotentHint=false, but the description adds real value: it explains 'detached' behavior, the immediate job-id return, the default user-session environment, and the >15s duration threshold. It doesn't mention how long jobs persist (though timeout param covers it) or rate limits, but this is solid additional context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences: first states what the tool does, second states when to use it. Zero filler, front-loaded action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a job-starting tool with no output schema: it explains the job-id return, the detachment, the environment, and the use-case threshold. Missing a note on how to check job status (sailfish_device_command_status) or cancel, but that's a routing detail rather than a correctness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are documented in the schema with rich detail (USB alias semantics, argv-not-shell, lifetime, runuser). The description adds only the 'user-session environment' framing, which is useful but doesn't go beyond what the schema provides for individual parameters. Baseline 3 for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (start), a specific resource (detached SSH command job), and its key mechanism (returns a job id). Unambiguously distinguished from siblings like sailfish_device_user_session_command and sailfish_device_command_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names use cases (GDB, tracing) and a threshold (>15s), which is clear routing guidance. It doesn't explicitly name the alternative sibling (e.g., a non-detached synchronous command tool), so it's a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_command_statusDevice Command Job StatusARead-onlyIdempotent
List detached device command jobs or read one job's status. A bounded failure excerpt is included automatically; running logs are omitted by default.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | No | Trailing log lines to include. Keep zero while the job is running; failed jobs include a bounded failure excerpt. | |
| job_id | No | Job id returned by sailfish_device_command_start. Omit to list recent jobs. | |
| verbose | No | Include full command, stdout, and stderr details for diagnostics. | |
| max_bytes | No | ||
| log_offset | No | Read a bounded log page from this byte offset. | |
| wait_seconds | No | Wait at most 15 seconds for a terminal state. | |
| after_revision | No | Last revision; unchanged results omit repeated details. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine behavioral context the annotations lack: a bounded failure excerpt is included automatically and running logs are omitted by default, telling the agent what to expect from the payload.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler, front-loading the dual capability and then the payload behavior. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter read tool with no output schema, the description covers the two modes and the log/failure-excerpt return behavior, which is the key thing an agent needs. It stops short of describing pagination or revision semantics, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so parameters like lines, job_id, verbose, and wait_seconds are already documented in the schema. The description restates the failure-excerpt/log behavior rather than adding new parameter-level syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List detached device command jobs or read one job's status') and clearly identifies the two operating modes. It distinguishes itself from the command lifecycle siblings by linking to sailfish_device_command_start, though it does not explicitly name which sibling to use for cancellation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The listing versus single-job modes are implied by 'Omit to list recent jobs' in the job_id parameter, but the description itself gives no explicit when-to-use framing such as polling a started job or choosing this over sailfish_device_command_cancel. Usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_install_rpmInstall Device RPMB
Start a detached RPM copy/install transaction; use sailfish_build_status/cancel.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Legacy synchronous execution; default returns a job for sailfish_build_status/cancel. | |
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | Command timeout in seconds. | |
| rpm_path | No | Single local RPM path. Use rpm_paths for dependency sets. | |
| installer | No | pkcon | |
| rpm_paths | No | Local RPM paths to copy and install in one transaction. | |
| remote_dir | No | Remote directory for copied RPMs. Created when missing. | |
| remote_path | No | Remote file path override for a single rpm_path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds the genuine behavioral fact that the operation is detached and hands back a job, which the annotations do not convey. It omits other behavioral traits the agent would want: files are copied to the device (side effects outside the RPM transaction), remote_dir is auto-created, and what happens to a partially-installed dependency set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words, so it is structurally efficient. It is arguably over-terse for an 8-parameter mutating tool, and the trailing 'use sailfish_build_status/cancel' reads like a boilerplate artifact rather than deliberate content, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, zero-required, mutating device operation with no output schema, the description covers only the async job handoff. It says nothing about failure modes, partial installs, what the job result contains, or that RPMs are physically copied to the target, leaving the agent under-informed about consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 88%, so the schema already documents device, wait, timeout, rpm_path vs rpm_paths, remote_dir, and remote_path with meaningful guidance (including the 'usb' alias convention). The description adds no parameter-level meaning beyond what the schema states, which is the baseline-3 case when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination ('Start a detached RPM copy/install transaction') that distinguishes it from siblings like sailfish_build_rpm and sailfish_device_command_start. However, the routing pointer to sailfish_build_status/cancel is confusing for a device-install tool (that naming pattern belongs to the build family), which slightly muddies rather than sharpens the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'detached' implies the asynchronous usage mode and the follow-up tools (status/cancel) are named, so an agent can infer the workflow. But there is no explicit when-to-use vs when-not guidance, no note on when to prefer the synchronous 'wait' path, and no prerequisite/permission guidance for a tool that mutates a device.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_journalRead Device JournalCRead-onlyIdempotent
Read recent journalctl output from a Sailfish OS device.
| Name | Required | Description | Default |
|---|---|---|---|
| grep | No | ||
| unit | No | ||
| lines | No | ||
| since | No | ||
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | Command timeout in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description only restates the read operation and adds no extra behavioral context such as output shape, streaming/pagination behavior, or what happens when the device is unreachable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, front-loaded with the essential verb and resource, with no filler. It is appropriately sized but arguably too terse for a six-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six parameters, only one required and low schema coverage, plus no output schema, the description is too thin. An agent gets no sense of return format, default windowing ('recent' is unquantified), or how grep/unit/since interact with the defaults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (device and timeout are documented, grep/unit/lines/since are not), and the description supplies no parameter meaning at all. It doesn't explain grep semantics, unit filtering, line limits, or the since time format, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource: reads journalctl output from a Sailfish OS device. It's clear what the tool does and reads distinct from siblings like sailfish_device_command_start or sailfish_device_topmost_pid. It doesn't explicitly differentiate itself from other log-adjacent siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives. The user might equally well run journalctl via sailfish_device_command_start or sailfish_device_user_session_command, and nothing tells the agent why to choose this one. No prerequisites or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_lipstick_screenshotLipstick ScreenshotB
Ask Lipstick to save a screenshot on the device, optionally pulling it locally.
| Name | Required | Description | Default |
|---|---|---|---|
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | Command timeout in seconds. | |
| local_path | No | ||
| privileged | No | ||
| remote_path | No | Lipstick accepts screenshot paths under the user home directory. | /home/defaultuser/Pictures/Screenshots/lipstick-<timestamp>.png |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds the meaningful behavioral nuance that it both saves on-device and can pull locally, but omits what happens on repeat calls, whether files overwrite, and permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and the optional local pull second. No wasted words, though it is very terse given the tool's five parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema and only 60% schema coverage, the description is adequate but thin: it does not explain what the tool returns (e.g., a path or handle), what privileged controls, or how the on-device path default interacts with local_path.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 60%; device, timeout and remote_path are documented in the schema, but local_path has no schema description and is only implied by 'pulling it locally'. The description adds marginal meaning for local_path but leaves privileged and timeout behavior unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: asking Lipstick (the Sailfish compositor) to save a screenshot on the device, with an optional local pull. This is clearly distinct from the surrounding device tools (touch, journal, install_rpm), though it doesn't explicitly name a sibling it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Optionally pulling it locally' implies when to set local_path, but there is no guidance on when to use this versus other device capture tools, no prerequisites (running session, device setup), and no explanation of when privileged or a non-default timeout is warranted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_proc_mapsRead Process MapsCRead-onlyIdempotent
Read or filter /proc//maps on a Sailfish OS device.
| Name | Required | Description | Default |
|---|---|---|---|
| pid | Yes | ||
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | Command timeout in seconds. | |
| contains | No | ||
| max_lines | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered. The description adds only that the underlying resource is a proc maps file; it discloses nothing about permission requirements, filter behavior, or output size limits beyond what the schema defaults imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the verb and resource front-loaded and no wasted words. It is appropriately sized, though its brevity is partly why parameter semantics suffer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A five-parameter tool with 40% schema coverage and no output schema should explain the filtering parameters and default behavior; the description covers none of that. It is too thin for the tool's surface area, even accounting for the rich annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40%, so the description must compensate and does not. The word 'filter' hints at 'contains' and 'max_lines', but neither is explained, and 'pid' has no description at all despite being the only required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (read/filter) and a precise resource (/proc/<pid>/maps) on a Sailfish OS device. An agent can tell this apart from sibling tools like sailfish_device_journal or sailfish_device_topmost_pid, though the description never names an alternative to disambiguate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to reach for this tool versus alternatives, nor any prerequisites (e.g., requires a running device or root access to inspect other processes' maps). Usage must be inferred entirely from the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_restart_serviceManage Device ServiceC
Run systemctl start/stop/restart/status for a system or user service.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | system | |
| unit | Yes | ||
| action | No | restart | |
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | Command timeout in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is largely covered. The description adds nothing beyond that — it does not warn that stop/restart disrupts a running service, nor mention permissions, device connectivity requirements, or what 'status' returns. Only the passing mention of 'system or user service' adds context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tightly written sentence that leads with the underlying command, so structure is good. It is arguably too terse for a 5-parameter mutation-capable tool, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, 40% schema coverage, no output schema, and no annotations covering permissions or side effects, the one-line description leaves substantial gaps: which device targets are valid, what happens on stop, and how the action/mode interact are all unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40% across 5 parameters, so the description must compensate and does not. The only parameters it touches are the actions (restart/start/stop/status) and the system/user distinction, which are already enumerated in the schema; unit, device, and timeout receive no additional explanation in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete verb (run) and resource (systemctl service control), and the action list implicitly tells the agent this is a process-control tool rather than a build or observation tool. It does not name or distinguish itself from any sibling, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to choose this tool over alternatives such as sailfish_device_command_start, sailfish_device_command_status, or sailfish_device_user_session_command, all of which could plausibly cover similar service-related needs. No prerequisites or preconditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_devicesList Sailfish DevicesARead-onlyIdempotent
List configured Sailfish OS device aliases and SSH targets.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered elsewhere. The description adds only the scoping detail that results are drawn from configured aliases and SSH targets — modest extra context about what 'device' means here, but no disclosure of ordering, environment sources, or failure behavior when nothing is configured.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, hedging, or redundant restatement of the title. Every word contributes scope information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with rich annotations and no output schema, this is nearly complete; the only gap is that nothing tells the agent how devices are enumerated or what happens when the configuration is empty. That gap is small given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4: there is nothing for the description to disambiguate. It correctly stays silent rather than inventing parameter behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and a specific resource (configured Sailfish OS device aliases and SSH targets), which is more precise than the generic 'devices'. It reads as a discovery/read tool distinct from the many sailfish_device_* action tools, though it never names a sibling to sharpen the boundary (e.g. vs. sailfish_device_setup, which likely also touches device configuration).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives, nor any prerequisite such as needing devices to be configured first. An agent can infer it is a discovery step before the device_* tools, but the description does not say so explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_setupCheck or Set Up Device PrerequisitesA
Check command availability and session configuration over SSH. By default this only reads device state. Explicit install_missing=true installs missing python3-base or mce-tools from existing repositories in a detached job; use sailfish_build_status/cancel. Missing base OS tools need manual repair.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | Wait for installation; use the default detached job for long work. | |
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | ||
| features | No | ||
| install_missing | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds context the annotations don't: the default invocation is read-only even though readOnlyHint is false, and installation only happens with an explicit flag, from existing repositories, in a detached job. That conditional behavior is exactly the kind of disclosure annotations cannot express. It omits what the detached job produces or how failure surfaces.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the default behavior, then the opt-in install behavior, then the limitation. No filler. Slightly dense and the pointer to 'sailfish_build_status/cancel' is ambiguous given the device-scoped siblings.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool at 40% schema coverage with no output schema, the description leaves real gaps: what 'features' selects, what a detached job returns, and how to distinguish a check failure from an unavailable command. It covers the safety-relevant install path well but not the full surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, so the description must compensate, and it does partly: install_missing's effect (installs python3-base or mce-tools from existing repos) is explained, and the detached-vs-wait distinction is implied. However, 'features' and 'timeout' are never addressed in the description, and features is a meaningful selector for the prerequisite check.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Check command availability and session configuration over SSH.' The default read vs. explicit install split is clear. It doesn't explicitly contrast itself with the closest sibling (sailfish_doctor), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating conditions: default is read-only, install_missing=true installs packages, missing base OS tools must be repaired manually. It even routes job tracking to build/command status and cancel tools. It stops short of stating when to pick this over sailfish_doctor or sailfish_device_install_rpm.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_topmost_pidTopmost Window PIDBRead-onlyIdempotent
Query Lipstick for the current topmost window process id.
| Name | Required | Description | Default |
|---|---|---|---|
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| timeout | No | Command timeout in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is well covered. The description adds only that the query goes to Lipstick; it does not state the return shape, latency, or what happens when no window is focused, which matters for an openWorld read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the key noun phrase (topmost window process id) front-loaded and zero filler. It is appropriately sized, though it is terse enough that more routing context could have been added without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-arg, read-only query with a fully documented schema and annotations, the description is minimally adequate. With no output schema, it only implicitly signals the return value (a pid) and omits any note about the no-focused-window case, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema thoroughly documents both the device alias/USB syntax and the timeout. The description contributes nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Query) and resource (topmost window process id) with the source system (Lipstick) named. It is easy to distinguish from device_journal or proc_maps, though the description does not explicitly contrast itself with any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as sailfish_device_proc_maps or sailfish_device_lipstick_screenshot, nor any prerequisite or exclusion stated. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_touchDevice Touch InputC
Discover the touchscreen input device or inject tap/swipe events over SSH.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| end_x | No | ||
| end_y | No | ||
| steps | No | ||
| action | Yes | discover lists input devices; tap and swipe inject Linux input events. | |
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| hold_ms | No | ||
| start_x | No | ||
| start_y | No | ||
| timeout | No | Command timeout in seconds. | |
| duration_ms | No | ||
| input_device | No | Optional explicit device path, for example /dev/input/event5. | |
| include_evdev_trace | No | Also run evdev_trace -i during discovery. Disabled by default because it may block on some devices. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the tool is not read-only, is open-world, non-idempotent, and non-destructive. The description adds that interaction occurs over SSH and that tap/swipe inject events, which is modest additional context but does not cover side effects, device requirements, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler, cleanly presenting the two main capabilities. It is appropriately concise, though its brevity contributes to the larger completeness gap.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's high complexity (14 parameters), low schema description coverage, and no output schema, the one-sentence description is far too sparse. It omits parameter semantics, usage guidance, side effects, and device/prerequisite details that an agent would need to invoke it accurately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low at 36%, and the description does not compensate: it repeats the action categories already covered by the enum description while saying nothing about the many coordinate, timing, or device-path parameters. For a 14-parameter tool, this leaves substantial ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states specific actions (discover, tap, swipe) and the target resource (touchscreen input device) over SSH. It clearly distinguishes the two operational modes, though it does not differentiate this tool from the sibling sailfish_device_touch_workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that discover and tap/swipe are alternative modes, but it gives no explicit guidance on when to use this tool versus the touch workflow sibling or when-not to use it. No prerequisites or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_touch_workflowScreenshot And TouchC
Capture a Lipstick screenshot, optionally list touch inputs, then inject a tap or swipe.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| end_x | No | ||
| end_y | No | ||
| steps | No | ||
| action | Yes | ||
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| hold_ms | No | ||
| start_x | No | ||
| start_y | No | ||
| timeout | No | Command timeout in seconds. | |
| local_path | No | Optional local path for the before screenshot. | |
| privileged | No | ||
| duration_ms | No | ||
| input_device | No | Optional explicit device path, for example /dev/input/event5. | |
| discover_input | No | ||
| after_local_path | No | ||
| screenshot_after | No | ||
| after_remote_path | No | ||
| before_local_path | No | ||
| screenshot_before | No | ||
| before_remote_path | No | ||
| include_evdev_trace | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds little beyond restating the steps: it does not disclose that input injection is a side effect on a live device, the implications of the privileged default, ordering of before/after screenshots, or any auth/rate constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler waste. It is efficient, though it trades away almost all useful detail for that brevity given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 23-parameter, multi-step device tool with 17% schema coverage and no output schema, a one-sentence description is far too thin. The workflow ordering, coordinate semantics, and screenshot/trace behavior that an agent needs to call it correctly are not conveyed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 23 parameters and only 17% schema description coverage, the description carries nearly the full burden and fails to help: it never explains x/y/start_x/end_x coordinates, steps, hold_ms, duration_ms, timeout, or the screenshot/trace toggles. Only 'tap or swipe' loosely maps to the action enum, leaving most parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names specific capabilities (capture a Lipstick screenshot, list touch inputs, inject a tap or swipe) with clear verbs and resources. However, it never distinguishes this bundled 'workflow' from its atomic siblings sailfish_device_touch and sailfish_device_lipstick_screenshot, so an agent cannot tell when the combined tool is preferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. The name 'workflow' hints that it bundles steps, but the description does not explain preferring this over the individual screenshot and touch tools, nor any prerequisites or device/connection conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_user_bus_callUser Bus CallC
Run a typed dbus-send method call on the defaultuser session bus.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| member | Yes | ||
| timeout | No | Command timeout in seconds. | |
| arguments | No | Raw dbus-send argument strings, for example string:/tmp/file. | |
| interface | Yes | ||
| destination | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, and destructiveHint=false, so the safety profile is largely covered structurally. The description adds little beyond that — it names the 'defaultuser session bus' scope but says nothing about mutation risk, required permissions, side effects, or response behavior for a tool that can invoke arbitrary dbus methods.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with no filler, and the core action is front-loaded. It is concise to the point of under-specification, but it wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, no-output-schema tool that can invoke arbitrary dbus methods on a session bus, the description is far too thin. It omits parameter meaning, usage conditions, and any behavioral/return context, leaving substantial gaps unfilled by schema or annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43% across 7 parameters, and the description contributes no parameter meaning at all. It never explains destination, path, interface, member, or arguments, leaving the required dbus-send addressing fields undocumented outside the schema while the coverage gap remains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Run') and resource ('typed dbus-send method call on the defaultuser session bus'), which is more than a tautology. However, it does not distinguish this tool from the adjacent sibling sailfish_device_user_session_command, leaving the agent to guess which session-level tool to use.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no when-not-to-use, and no named alternative. The description merely states what the tool does, not the situation that should select it over sailfish_device_user_session_command or sailfish_device_command_start.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_device_user_session_commandUser Session CommandA
Run a short command, limited to 15 seconds, with the configured Sailfish user-session D-Bus environment. Use sailfish_device_command_start for anything that may run longer.
| Name | Required | Description | Default |
|---|---|---|---|
| device | No | Configured device alias or ssh target. When the user says the device is attached over USB, use 'usb' for the default device or 'usb:<name>' when the user supplies a device name. The USB connection trusts that name for device metadata and does not resolve it through DNS. | |
| command | Yes | Command argv to run; shell parsing is not applied. | |
| timeout | No | Command timeout in seconds; long commands must use an async job. | |
| verbose | No | Include full command, stdout, and stderr details for diagnostics. | |
| run_as_user | No | Run through runuser/su as the configured device username. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readonly, open-world, non-idempotent, and non-destructive behavior. The description adds meaningful execution context beyond annotations by specifying the 15-second limit and the configured user-session D-Bus environment, though it does not discuss permissions or return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted wording. The key constraint and the alternative tool are both front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a command-execution tool with full schema coverage and annotations covering safety, the description is nearly complete for selection and invocation. It does not describe result/output behavior, which leaves a minor gap because there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter semantics are already well documented in the schema. The description does not add parameter-level syntax or format details beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: run a short command in the configured Sailfish user-session D-Bus environment. It also distinguishes itself from sailfish_device_command_start, so an agent can tell them apart without opening both schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the alternative for longer-running commands: use sailfish_device_command_start for anything that may run longer. The short-command scope is front-loaded, making the when-to-use boundary clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_doctorARead-onlyIdempotent
Check helper compatibility, local tools, SDK targets and configuration without contacting devices.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and open-world behavior. The description adds a genuinely useful behavioral constraint not in the annotations: it does not contact devices, which tells the agent the check is safe and side-effect-free on hardware. It does not describe output shape, but the safety profile is otherwise covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The key constraint ('without contacting devices') is placed at the end of the clause but the enumeration is compact and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic with rich annotations and no output schema, the description adequately conveys scope. It could note what the check reports (pass/fail, missing helpers) but nothing required for correct invocation is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty with 100% coverage, so there is no parameter semantics burden. Baseline 4 applies since nothing is missing or contradictory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Check) plus an enumerated set of resources (helper compatibility, local tools, SDK targets, configuration). This clearly separates it from the device-facing siblings. It stops short of naming a sibling alternative, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without contacting devices' implies the context of use (local/offline diagnostics as opposed to the many sailfish_device_* tools), but there is no explicit when-to-use or when-not-to-use statement. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_obs_buildlogOBS Build LogCRead-onlyIdempotent
Fetch a build log from a configured OBS server or explicit API alias; defaults to the API nostream form.
| Name | Required | Description | Default |
|---|---|---|---|
| arch | Yes | ||
| lines | No | Trailing log lines to return. Use zero with verbose=true for the legacy full-output diagnostic response. | |
| server | No | Named OBS server configured in paths.obs_servers. | |
| package | Yes | ||
| project | Yes | ||
| timeout | No | Command timeout in seconds. | |
| verbose | No | Include full command, stdout, and stderr details for diagnostics. | |
| nostream | No | ||
| api_alias | No | Advanced raw osc -A alias or API URL; overrides paths.osc_api_alias and cannot be combined with server. | |
| repository | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and non-destructive, so the safety profile is covered. The description adds the default-transport behavior (nostream form) and the server-vs-alias sourcing model, but says nothing about truncation, pagination, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler; the source options and default are stated compactly. It is efficient, though very short given the tool's parameter surface.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With ten parameters, only half documented, no output schema, and the required identifiers unexplained, the description is too thin for the tool's complexity. An agent has little basis for choosing between server and api_alias or for interpreting the returned log.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%, and the four required parameters (project, package, repository, arch) are undocumented in both schema and description. The description only hints at the server/api_alias relationship, leaving most of the ten parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (build log) and names the source options (configured OBS server vs explicit API alias). It does not differentiate itself from close siblings like sailfish_obs_results or sailfish_obs_watch, which an agent must choose between.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives no when-to-use guidance, no prerequisites, and no exclusions relative to the many sibling tools. The only guidance offered is an implementation default ('defaults to the API nostream form'), which is not selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_obs_resultsOBS ResultsARead-onlyIdempotent
Read OBS results. Use wait_seconds for bounded mechanical monitoring instead of repeated model-driven status calls.
| Name | Required | Description | Default |
|---|---|---|---|
| server | No | Named OBS server configured in paths.obs_servers. | |
| package | No | ||
| project | Yes | ||
| timeout | No | Command timeout in seconds. | |
| verbose | No | Include full command, stdout, and stderr details for diagnostics. | |
| api_alias | No | Advanced raw osc -A alias or API URL; overrides paths.osc_api_alias and cannot be combined with server. | |
| wait_seconds | No | Mechanically wait at most 15 seconds while results remain pending. | |
| poll_interval | No | Seconds between OBS checks while waiting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is largely covered. The description adds that waiting is bounded and mechanical rather than model-driven, which is useful behavioral context. It still omits return shape, pending-state semantics, and auth/diagnostic behavior that a read tool without an output schema should ideally disclose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the purpose and followed by a specific usage directive. Every sentence carries weight, and there is no redundant restatement of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter read tool with no output schema, the description is minimal: it never describes what 'OBS results' contain or how to interpret pending/failure states. Annotations and 75% schema coverage fill some gaps, but the missing return-value guidance and sibling differentiation leave real holes. Adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so most parameters are documented in the schema itself. The description only elaborates on wait_seconds, adding rationale beyond the schema's 'mechanically wait' wording, but it says nothing about the undocumented project and package parameters or interactions among server, api_alias, and timeout. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Read OBS results states a clear verb and resource, so an agent can tell this is a read operation on OBS output. However, 'results' is underspecified and the description never differentiates this tool from close siblings like sailfish_obs_watch or sailfish_obs_buildlog. Clear purpose, but no sibling routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides one explicit usage rule: use wait_seconds for bounded mechanical monitoring instead of repeated model-driven status calls. That is useful guidance for avoiding polling loops, but it never states when to choose this tool over sailfish_obs_watch or other OBS/build-status siblings, and it offers no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_obs_watchB
Watch OBS until completion in a detached local job; use sailfish_build_status/cancel or the jobs CLI waiter.
| Name | Required | Description | Default |
|---|---|---|---|
| server | No | Named OBS server configured in paths.obs_servers. | |
| package | No | ||
| project | Yes | ||
| timeout | No | Command timeout in seconds. | |
| api_alias | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, so the safety profile is largely covered. The description adds genuinely useful context the annotations don't carry: this blocks until completion and runs as a detached local job. However, it omits what happens on timeout, whether the job is retrievable afterward, and whether credentials to the OBS server are required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence that front-loads the core behavior before the alternative-tool pointer. It is short and has no filler, though the trailing clause reads as an incomplete sentence and slightly muddies the routing advice.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, non-idempotent, open-world tool with no output schema and only 40% schema coverage, the description is too thin. It doesn't explain the required project param, the OBS package/project targeting, api_alias, or what the caller gets back from a detached watch job.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% (server and timeout are documented; package, project, and api_alias are not), and the description adds no parameter meaning at all. It never mentions project (required), package, or api_alias, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Watch OBS until completion') plus the execution mechanism ('detached local job'), which is more specific than the title alone. It orients the agent against siblings by naming sailfish_build_status/cancel as the polling alternatives, though the distinction is compressed into a clause.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gestures at alternatives (sailfish_build_status/cancel, jobs CLI waiter) but as a bare fragment, leaving it unclear whether these are substitutes for this tool or follow-up tools. There is no explicit when-to-use/when-not statement, no prerequisites, and no statement of what 'detached' means for the caller.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_qml_check_translator_ternariesCheck QML Ternary TranslationsCRead-onlyIdempotent
Flag QML ternary qsTrId expressions that need branch-local comments.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds the narrow detection rule being enforced but says nothing about output shape (errors vs. list of findings), exit behavior, or whether a path defaults; with annotations present, this lands at the baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the detection target stated immediately and no filler. It is tight, though the brevity comes at the cost of the missing usage and parameter detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, no documentation of the only parameter, and no statement of what a run produces or requires. For a checker whose results an agent must interpret, the definition leaves too much unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the lone 'path' parameter at all — whether it is a file, directory, or project root, or what the default is when omitted. Since the schema carries no semantic detail, the description had to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Flag') and a precise resource ('QML ternary qsTrId expressions that need branch-local comments'), so the agent knows exactly what lint class this covers. It does not, however, distinguish itself from the closely related sibling sailfish_qml_find_translations, so differentiation is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to reach for this checker versus sailfish_qml_find_translations, no mention of prerequisites (e.g., a valid QML project root), and no scope/exclusion guidance. Usage must be inferred entirely from the one-line purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_qml_find_translationsFind QML TranslationsCRead-onlyIdempotent
Find qsTrId and translator comments in QML files.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| timeout | No | Command timeout in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint, and openWorldHint, so the safety profile is covered. However, the description adds nothing beyond stating what is searched – no mention of scope (single file vs directory tree), no output format, no behavioral context about the 'path' parameter's effect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, efficient sentence with no filler. It is front-loaded and readable, though extremely terse given the tool's search behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with an undocumented search-scope parameter and no output schema, the description leaves significant gaps: what 'path' means (file vs directory), whether the search is recursive, and what the result looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%; 'timeout' is documented in the schema while 'path' has no description anywhere. The description does not compensate for the undocumented 'path' parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Find) and resource (qsTrId and translator comments in QML files), which is clear and distinguishable from most siblings. It doesn't explicitly differentiate from sailfish_qml_check_translator_ternaries, the closest sibling, but the scope is specific enough to select the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of the alternative sibling sailfish_qml_check_translator_ternaries. The agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_repo_findRepo FindCRead-onlyIdempotent
Search a repo or subtree under the configured git root.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| query | Yes | ||
| timeout | No | Command timeout in seconds. | |
| max_count | No | ||
| fixed_strings | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds the scope constraint (configured git root) but says nothing about result format, ordering, truncation at max_count, or timeout behavior, so its added value is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler, which is structurally clean. But the brevity comes at the cost of under-specification rather than being genuinely concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with 20% schema coverage and no output schema, the single-sentence description leaves key calling details (query syntax, path scoping, result limits) undocumented. Annotations cover the safety profile, but the operational contract is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (only timeout is documented), yet the description adds no meaning for query, path, max_count, or fixed_strings. It does not clarify whether query is a regex or literal by default, what path scoping means, or how max_count truncation behaves, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (repo/subtree under the configured git root), which is clearer than a bare name restatement. However, it does not differentiate from siblings like sailfish_repo_status or sailfish_qml_find_translations, and 'find' versus 'search' semantics (pattern matching? grep?) are left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternatives are given. The only guidance is implicit in the phrase 'under the configured git root', which narrows scope but doesn't route the agent between this tool and related repo/qml search siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_repo_statusRepo StatusBRead-onlyIdempotent
Run git status --short --branch under the configured git root.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| timeout | No | Command timeout in seconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/openWorld, so the safety profile is covered. The description adds real value by naming the exact invocation, which implicitly reveals the output shape (short status with branch). It omits failure behavior (e.g. non-git directory) but is strong for a read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler that conveys the verb, flags, and scope. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, naming the exact git invocation largely documents the return. But the undocumented 'path' parameter and lack of any error/edge-case note leave gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: 'timeout' is documented, but 'path' has no schema description and the description does not explain it. The phrase 'under the configured git root' only hints at path semantics and does not compensate for the undocumented parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and the exact command executed ('Run git status --short --branch'), so the agent knows precisely what the tool does. However it does not differentiate itself from the sibling sailfish_repo_find, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus sailfish_repo_find or other repo/spec tools, and no prerequisites or exclusions stated. The agent must infer usage from the command name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_sdk_refresh_metadataRefresh SDK MetadataB
Start a detached SDK metadata refresh; use sailfish_build_status/cancel.
| Name | Required | Description | Default |
|---|---|---|---|
| arch | No | ||
| wait | No | Legacy synchronous execution; default returns a job for sailfish_build_status/cancel. | |
| force | No | Append zypper ref -f. | |
| device | No | Optional configured device to supply default release and architecture. | |
| target | No | Local SDK target base or .default target, for example aarch64 or aarch64.default. | |
| release | No | ||
| timeout | No | Command timeout in seconds. | |
| local_sdk | No | Override paths.local_sdk for this call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false, openWorldHint=true), so the safety bar is partly met. The description adds the meaningful 'detached' detail (returns a job rather than blocking), but omits auth requirements, what metadata is actually rewritten, and whether an existing refresh is superseded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, fully front-loaded with the action and the routing hint. It earns its place, though it is terse enough that a little more context would not have hurt.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 optional params, a mutation profile, and no output schema, the description covers the essential async pattern but leaves gaps: no mention of what the returned job is called, how arch/release/target interact, or the side effects of force. Adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, with the notable params (wait, force, device, target, timeout, local_sdk) documented in-schema. The description adds no parameter meaning beyond that, so the baseline 3 applies for a schema that largely carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Start') and resource ('SDK metadata refresh'), and signals the execution model with 'detached'. It partially differentiates from siblings by pointing at sailfish_build_status/cancel, though it does not distinguish itself from adjacent tools like sailfish_repo_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use sailfish_build_status/cancel' implies the async job workflow and implies when to check status, but there is no explicit when/when-not guidance versus alternatives such as sailfish_repo_status or sailfish_build_rpm. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sailfish_spec_summaryRPM Spec SummaryCRead-onlyIdempotent
Parse high-level metadata from a Sailfish RPM spec file.
| Name | Required | Description | Default |
|---|---|---|---|
| repo_path | No | ||
| spec_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, open-world behavior, so the safety profile is covered. The description adds nothing beyond that — it doesn't say what happens if the spec is missing/malformed, whether it resolves includes/macros, or what form the parsed result takes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and no filler. Efficient, though it is arguably too terse for a tool with two undocumented parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining what 'high-level metadata' is returned (name/version/release/license?), and it doesn't. Combined with two undescribed parameters and no guidance on repo_path vs spec_path, the definition leaves real gaps for an agent trying to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and neither repo_path nor spec_path is mentioned in the description, so the agent gets no help understanding the two paths or what happens when both (or neither, since neither is required) are supplied. The phrase 'from a ... spec file' weakly implies spec_path but leaves repo_path wholly unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (parse) and resource (high-level metadata from a Sailfish RPM spec file), which cleanly separates it from build/repo/device siblings. It stops short of naming what 'high-level metadata' actually contains or which sibling would be used instead, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to reach for this tool versus the many build, OBS, and repo siblings that also touch packaging data. No prerequisites, no mention of when a spec file should be inspected at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
35 tool updates
v0.1.0- First observed
sailfish_android_build - First observed
sailfish_android_build_cancel - First observed
sailfish_android_build_hosts - First observed
sailfish_android_build_status - First observed
sailfish_build_cancel - First observed
sailfish_build_preflight - First observed
sailfish_build_rpm - First observed
sailfish_build_status - First observed
sailfish_device_app_launch - First observed
sailfish_device_browser_launch - First observed
sailfish_device_command_cancel - First observed
sailfish_device_command_start - First observed
sailfish_device_command_status - First observed
sailfish_device_install_rpm - First observed
sailfish_device_journal - First observed
sailfish_device_lipstick_screenshot - First observed
sailfish_device_proc_maps - First observed
sailfish_device_restart_service - First observed
sailfish_device_setup - First observed
sailfish_device_topmost_pid - First observed
sailfish_device_touch - First observed
sailfish_device_touch_workflow - First observed
sailfish_device_user_bus_call - First observed
sailfish_device_user_session_command - First observed
sailfish_devices - First observed
sailfish_doctor - First observed
sailfish_obs_buildlog - First observed
sailfish_obs_results - First observed
sailfish_obs_watch - First observed
sailfish_qml_check_translator_ternaries - First observed
sailfish_qml_find_translations - First observed
sailfish_repo_find - First observed
sailfish_repo_status - First observed
sailfish_sdk_refresh_metadata - First observed
sailfish_spec_summary
TDQS
Scored across 35 tools
Tools are grouped by domain and most have clearly distinct purposes, such as local RPM build, Android build, OBS, device SSH commands, and QML translation checks. Some boundaries blur because sailfish_build_status/cancel are reused for multiple local async jobs despite the 'build' name, and device_user_session_command vs device_command_start require reading descriptions to distinguish short from long-running commands.
All tool names use snake_case with a consistent sailfish_ prefix, which makes the set predictable. However, the suffix pattern is not uniformly verb_noun: sailfish_devices, sailfish_doctor, sailfish_spec_summary, and sailfish_obs_watch mix noun/state and action forms, though deviations are minor.
35 tools is heavy for a single MCP server, even though the server spans several subdomains including device, build, Android, OBS, repo, SDK, spec, and QML. The breadth justifies many tools, but the surface is at the upper end and could likely be consolidated in places.
The tool set covers a wide lifecycle: build planning, local and Android builds, device inspection and command execution, OBS monitoring, SDK refresh, repo search, spec parsing, and QML translation checks. Minor gaps remain, such as RPM uninstall, explicit app stop/termination, git write operations, and signing/deployment workflows.
Maintenance
Related MCP Connectors
- TypeshipOAuthdev.typeship
Generate a typed SDK, CLI, and MCP server from any OpenAPI or GraphQL spec, and keep them current.
MCP access to ELSHWORK agents, repositories, isolated runs, discovery and tasks.
MCP tools for TON Sites, TON DNS and TON Storage.
Manage CloudPepper servers, Odoo instances, backups, and deployments over MCP.
Related MCP Servers
- FlicenseDqualityBmaintenanceEnables HarmonyOS device discovery, app build and deployment, UI automation, E2E inspection, and log validation through MCP tools.1812-
- AlicenseAqualityFmaintenanceEnables building, verifying, previewing, deploying, diagnosing, updating, and rolling back full-stack Micro apps through typed local MCP tools backed by micro-cli.6828 npmMIT
- FlicenseNot gradedqualityCmaintenanceEnables agents to connect to remote MCP servers once, access their tools through a compact MCP endpoint, pair a CLI inside sandboxes, and create watches that turn command or tool output into pollable structured events.2-
- FlicenseNot gradedqualityCmaintenanceEnables secure discovery and invocation of sandboxed filesystem, repository inspection, and utility tools through a unified MCP client with schema validation, timeouts, and execution traces.-