Skip to main content
Glama

An MCP server to query, troubleshoot and manage Sonarr, Radarr and Prowlarr from Claude Desktop or any other MCP client.

It focuses on the questions that come up most often in the Sonarr and Radarr issues and wiki:

  • "Why didn't it grab S03E04? Which releases are out there and why were they rejected?"

  • "Why isn't this download importing?"

  • "Which downloads are stalled? Remove them and find another release."

  • "Would this release be an upgrade over my file? Why does it say 'not an upgrade'?"

  • "Remove the .exe releases from the queue, blocklist them and search again."

  • "Which indexer failed the most this month?"

  • "Add Severance with the HD-1080p profile."

Status: v0.2.0. 43 tools, 269 tests. Validated against a real Sonarr, Radarr and Prowlarr on a Synology NAS, and every tool (including writes) against disposable Sonarr/Radarr/Prowlarr containers. See Status.

Features

  • 42 tools in seven groups: status, library, diagnostics, troubleshooting, maintenance, actions and Prowlarr. Full reference in docs/TOOLS.md.

  • Guided troubleshooting:

    • explain_search runs an interactive search and groups the releases by rejection reason (language, size, quality, custom format score…).

    • diagnose_import and check_paths give the probable cause of a failed import and how to fix it: remote path mapping, permissions, manual match, "not an upgrade"…

    • find_stalled_downloads spots stuck downloads.

    • parse_release shows how a release is scored and estimates whether it would be an upgrade.

    • get_health adds the usual fix to each known health warning; test_download_clients tests the download clients.

  • Maintenance: bulk_edit_preview / bulk_edit_apply change many series/movies at once (monitoring, quality profile, tags…), always with a preview first; preview_rename / rename_files rename with a preview; library_report shows what takes space and what could be cleaned up.

  • Follows Sonarr/Radarr good practice, enforced in code: an hourly indexer budget, no repeated searches, mass searches and removals in two steps, imports that never break seeding or replace files without permission, and previews before bulk changes. See docs/GOOD_PRACTICES.md.

  • Compact answers. Never returns raw API JSON: overviews are trimmed, images and internal fields are dropped, and lists are paged with limit and truncated.

  • Readable errors in English or Spanish (ARR_LANG). For example, "Sonarr returned 401: wrong API key (check SONARR_API_KEY)" instead of a traceback.

  • Titles or IDs. Tools accept "The Capture", "Dune 2021" or 42. When a title is ambiguous they return the candidates instead of guessing.

  • Partial failures are tolerated. If Radarr is down, get_queue still shows Sonarr's queue and flags Radarr's error.

  • Read-only mode (ARR_READONLY=true): write tools are not even registered.

  • Safe imports. manual_import only imports files listed one by one, never imports two files for the same episode, and imports nothing if anything is wrong. Unidentified files can be matched by hand; Sonarr/Radarr re-validate the match before importing.

  • Fake releases (.exe .scr .lnk .bat .cmd .msi .zipx) and orphaned downloads (files left in the download folder that nothing is tracking).

Related MCP server: media-stack-mcp

Requirements

  • Windows or Linux (tested in CI on both, Python 3.11–3.13; macOS should work but is not tested).

  • Python ≥ 3.11

  • Sonarr v4 and/or Radarr v5–v6 (/api/v3). Tested with Sonarr 4.0.16 and 4.0.20, Radarr 5.28, 6.2 and 6.4.4. Prowlarr (/api/v1, tested with 2.4.0 and 2.6.5) is optional.

  • The API key of each service (in each app: Settings → General → Security → API Key).

Installation

Quickest: with uv (no Git, no manual Python setup)

uv downloads Python by itself and runs arr-doctor straight from the latest release.

  1. Install uv once:

    • Windows (PowerShell): powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

    • Linux / macOS: curl -LsSf https://astral.sh/uv/install.sh | sh

  2. In Claude Desktop (Settings → Developer → Edit Config), add inside mcpServers:

    "arr": {
      "command": "uvx",
      "args": ["--from", "https://github.com/alexliam83/arr-doctor/releases/download/v0.2.0/arr_doctor-0.2.0-py3-none-any.whl", "arr-doctor"],
      "env": {
        "SONARR_URL": "http://192.168.1.10:8989",
        "SONARR_API_KEY": "your-sonarr-api-key",
        "RADARR_URL": "http://192.168.1.10:7878",
        "RADARR_API_KEY": "your-radarr-api-key"
      }
    }

    If Claude Desktop cannot find uvx, put its full path in command, e.g. "command": "C:\\Users\\<you>\\.local\\bin\\uvx.exe" on Windows or "~/.local/bin/uvx" elsewhere.

  3. Restart Claude Desktop completely.

Without uv: pip install https://github.com/alexliam83/arr-doctor/releases/download/v0.2.0/arr_doctor-0.2.0-py3-none-any.whl installs the arr-doctor command.

From source (Windows, Claude Desktop)

cd C:\
git clone https://github.com/alexliam83/arr-doctor.git
cd arr-doctor
python -m venv .venv
.\.venv\Scripts\pip install -e .

In Claude Desktop, open Settings → Developer → Edit Config and add this inside mcpServers, leaving the rest of the file alone:

"arr": {
  "command": "C:\\arr-doctor\\.venv\\Scripts\\arr-doctor.exe",
  "env": {
    "SONARR_URL": "http://192.168.1.10:8989",
    "SONARR_API_KEY": "your-sonarr-api-key",
    "RADARR_URL": "http://192.168.1.10:7878",
    "RADARR_API_KEY": "your-radarr-api-key"
  }
}

If the file was empty, it ends up like this:

{
  "mcpServers": {
    "arr": { "command": "…", "env": { "…": "…" } }
  }
}

Restart Claude Desktop completely, quitting it from the system tray icon too. In a new conversation, the tools icon should list arr.

Claude Desktop from the Microsoft Store. This version keeps its config in a virtualized path, usually %LOCALAPPDATA%\Packages\Claude_…\LocalCache\Roaming\Claude\claude_desktop_config.json. Always open it with "Edit Config" to be sure you edit the right file. Server logs are in the logs folder of that same directory (mcp-server-arr.log).

Linux / macOS

git clone https://github.com/alexliam83/arr-doctor.git && cd arr-doctor
python3 -m venv .venv
.venv/bin/pip install -e .
cp .env.example .env   # and fill in the keys

In your MCP client config, use /path/to/arr-doctor/.venv/bin/arr-doctor as the command.

Configuration

Variables are read from the environment (the env block of the MCP client) or from a .env file in the repository folder. There is a template in .env.example.

Variable

Required

Default

Description

SONARR_URL

yes*

—

Full URL, e.g. http://192.168.1.10:8989

SONARR_API_KEY

yes*

—

RADARR_URL

yes*

—

e.g. http://192.168.1.10:7878

RADARR_API_KEY

yes*

—

PROWLARR_URL

no

—

e.g. http://192.168.1.10:9696

PROWLARR_API_KEY

no

—

ARR_LANG

no

en

Language of errors and notes: en or es

ARR_READONLY

no

false

true disables every write tool

ARR_INDEXER_BUDGET_PER_HOUR

no

30

Max searches/grabs/indexer tests per hour (0 = unlimited)

ARR_SEARCH_CONFIRM_ABOVE

no

10

Searches or queue removals over this many items need confirmation

ARR_SEARCH_COOLDOWN_MINUTES

no

60

Skip items searched less than this many minutes ago (0 = off)

ARR_TIMEOUT

no

15

HTTP timeout in seconds. Scans and interactive searches use higher values

ARR_DOWNLOADS_PATH

no

auto

Download folder as Sonarr/Radarr see it, used by find_orphan_downloads. If unset, it is detected from the queue and the import history

* At least Sonarr or Radarr is required. A service without both a URL and an API key is not registered, so its tools do not appear. Full URLs are used so reverse proxies and HTTPS work.

Docker / NAS. If the port is assigned automatically (for example, Synology Container Manager in auto mode), it may change when the container is recreated. If get_system_status cannot connect, check the port, or better, pin it.

API keys never go into the repository. .env is in .gitignore.

Usage

Typical conversations and the tools Claude usually picks:

Question

Tools

"Why didn't it grab S03E04?"

explain_search → grab_release if you want a rejected one

"Why isn't it importing?"

diagnose_import → check_paths → manual_import(matches=…)

"Which downloads are stalled?"

find_stalled_downloads → remove_from_queue(blocklist=true)

"Would this release be an upgrade?"

parse_release

"What's stuck in the queue?"

get_queue(only_problems=true)

"Remove the .exe releases and search again"

find_suspicious_releases → remove_from_queue(blocklist=true)

"Which indexer fails the most this month?"

get_indexer_stats(days=30)

"What takes the most space?"

library_report

"Unmonitor every ended series that is complete"

bulk_edit_preview → bulk_edit_apply

"Which files don't follow my naming scheme?"

preview_rename → rename_files

"Add Severance in HD-1080p"

add_series (if profile or root folder are missing, it returns the options)

"Import what was left in downloads"

find_orphan_downloads → manual_import(files=[…])

Searches, refreshes and imports are asynchronous commands in Sonarr/Radarr. The tool returns a command_id you can follow with get_command_status.

explain_search queries every indexer live (it is the same interactive search as the web UI). Use it for a specific episode or movie, not in a loop.

Tools that change or delete things

  • remove_from_queue: removes queue items. By default it also removes the task from the download client (remove_from_client=true). With blocklist=true the release is blocklisted and, unless skip_redownload=true, Sonarr/Radarr automatically search for another one.

  • manual_import: imports specific files into the library. By default (import_mode="auto") it copies/hardlinks while the download is seeding and moves otherwise. Replacing a file already in the library needs allow_replace=true.

  • grab_release: sends a specific release to the download client even if it was rejected.

  • bulk_edit_apply: changes many items at once; for several items it needs the token from bulk_edit_preview.

  • rename_files: renames files on disk to match the naming settings.

Claude Desktop asks for permission the first time each tool is used. For these it is better not to choose "Always allow", so you can review each call. The two that delete also carry the MCP destructiveHint annotation.

Development

python -m venv .venv
.venv/bin/pip install -e ".[dev]"        # Windows: .venv\Scripts\pip
.venv/bin/pytest                          # unit tests with respx (no network)
uvx ruff check arr_doctor tests scripts      # lint
python scripts/gen_tools_doc.py           # regenerates docs/TOOLS.md and docs/TOOLS.es.md (a test checks they are up to date)

Against a real server (read-only, uses .env):

python scripts/smoke.py              # calls every read-only tool and prints a summary
python scripts/smoke.py --verbose    # also prints each answer
python scripts/smoke.py --capture    # saves anonymized raw responses to scripts/captures/ (git-ignored)

smoke.py refuses to call any tool that is not annotated readOnlyHint, and skips explain_search because it queries the indexers.

In Claude Desktop: a manual checklist (does the model pick the right tools, respect confirmations, explain results?) is in docs/testing/claude-desktop.md.

MCP Inspector:

npx @modelcontextprotocol/inspector .venv/bin/python -m arr_doctor.server          # web UI
npx @modelcontextprotocol/inspector --cli .venv/bin/arr-doctor -e SONARR_URL=… -e SONARR_API_KEY=… --method tools/list

On Windows, replace .venv/bin/python with .venv/Scripts/python.exe.

Layout

arr_doctor/
├── server.py         # FastMCP, conditional tool registration, main()
├── config.py         # Settings (pydantic-settings)
├── i18n.py           # English and Spanish messages
├── paths.py          # comparing POSIX / Windows / UNC paths from Sonarr/Radarr
├── formatting.py     # compacting API responses
├── clients/          # HTTP: base.py (errors, X-Api-Key), sonarr.py, radarr.py, prowlarr.py
└── tools/            # library, diagnostics, troubleshoot, maintenance, actions, indexers, common
tests/                # pytest + respx; fixtures/ with sample responses
scripts/              # smoke.py, gen_tools_doc.py
docs/                 # TOOLS (generated), DESIGN, ROADMAP — in English and Spanish

Status

All tools have unit tests (mocked HTTP) and the server has an end-to-end stdio test. On top of that, they were run against a real Sonarr 4.0.16 and Radarr 6.2 on a Synology NAS (real library, read-only plus a few approved changes) and against disposable Docker containers of Sonarr 4.0.20, Radarr 5.28 and 6.4.4, and Prowlarr 2.6.5 (empty libraries, every tool including writes):

Block

Real NAS

Containers

Notes

Skeleton, get_system_status

✅

✅

stdio, MCP Inspector, readable errors

Library and diagnostics (read-only)

✅

✅

Every tool, Sonarr and Radarr

Troubleshooting (read-only)

✅

✅

On the NAS it found a real fake .exe, stalled downloads and orphaned files. explain_search ran in containers only (no indexers there)

Write actions

✅ partly

✅

NAS: refresh_item, remove_from_queue with blocklist, manual_import (6 episodes). Containers: add_series/add_movie, set_monitored, all searches. grab_release only mocked (needs real indexers)

Maintenance

✅ partly

✅

NAS: library_report, preview_rename. Containers: bulk_edit_preview / bulk_edit_apply, test_download_clients. rename_files apply only mocked (needs files on disk)

Prowlarr

✅

✅

NAS (Prowlarr 2.4.0, 11 indexers): every tool, including test_indexers and set_indexer_enabled (8 dead indexers disabled; credentials untouched)

Windows and Linux

✅

—

CI on both + a real Windows machine

Good-practice guardrails

✅

✅

Roadmap low priority

—

—

Not started

Testing against the real server found and fixed several issues mocks could not show (unknown queue items, stalled-download cases, imports reported as successful that had failed, long imports vs. client timeouts). The fixtures in tests/fixtures/ still use made-up data shaped like the official OpenAPI schemas; they can be replaced with anonymized real responses (smoke.py --capture).

Credits and license

Created by Alexliam (@alexliam).

Inspired by BerryKuipers/mcp_services_radarr_sonarr. This is a new implementation that does not reuse its code.

MIT license.

Available Tools

36 tools
add_seriesA

Add a new series to Sonarr. Looks it up online first; if several results match it returns the candidates, and if quality profile or root folder are missing it returns the available options instead of guessing.

Examples: "Add The Capture with the HD-1080p profile", "Add Severance and search for it"

ParametersJSON Schema
NameRequiredDescriptionDefault
termYesTitle to look up online, or 'tvdb:12345'
monitorNoall
tvdb_idNoPick this exact result from the lookup
search_nowNoSearch for missing episodes right away
root_folderNoRoot folder path
quality_profileNoQuality profile name or ID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds meaningful behavior beyond the annotations: it discloses a lookup-first, two-phase flow and, crucially, that the tool returns candidates or available options rather than guessing when inputs are ambiguous or incomplete. Annotations already cover the safety profile (write, non-destructive), so this extra context is valuable. Gaps: no statement on idempotency (what happens if the series already exists) and the 'looks it up online' claim sits awkwardly against openWorldHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by the key behavioral caveats and two compact examples. No filler sentences; the examples earn their place by illustrating parameter usage. Slightly longer than strictly necessary but well organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers the main decision flow for a write tool. Remaining gap is handling of duplicates/already-existing series and any permission prerequisites, which an agent would want before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the baseline is 3, but the description adds real meaning: it clarifies that term is used for an online lookup and that quality_profile/root_folder may be absent, triggering an options response. The examples map natural-language phrases onto search_now and quality_profile, which the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add a new series to Sonarr') that an agent can immediately distinguish from read-oriented siblings like get_series or search_library. The purpose is unambiguous and reinforced by two concrete invocation examples.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the operational flow (look up online first; return candidates on ambiguity; return available options if quality profile or root folder are missing), which tells the agent what to expect and how to follow up. It does not explicitly name when to prefer this over any alternative, but no sibling competes for the 'add a series' job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bulk_edit_applyA

Apply a change to many series or movies at once. For several items or a filter selection, call bulk_edit_preview first, show it to the user and only call this once they agree, passing the preview's confirmation_token (same selection and changes). A single item named by the user can be changed directly. Root folders are not changed here because that would move files.

Examples: "Yes, unmonitor those ended series", "Tag The Capture with 'uk'"

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoFilter: has this tag
itemsNoSeries/movie IDs or titles to edit
statusNoFilter by status. Sonarr: continuing, ended, upcoming. Radarr: announced, inCinemas, released
serviceYes'sonarr' (series) or 'radarr' (movies)
year_toNoFilter: year <= this
add_tagsNoChange: tags to add (created if missing)
completeNoFilter: all episodes on disk (series) / has file (movies)
monitoredNoFilter: currently monitored or not
year_fromNoFilter: year >= this
remove_tagsNoChange: tags to remove
set_monitoredNoChange: monitor or unmonitor
quality_profileNoFilter: current quality profile name or ID
set_series_typeNoChange (Sonarr)
confirmation_tokenNoToken from bulk_edit_preview; required for several items or a filter selection
set_quality_profileNoChange: new quality profile name or ID
set_minimum_availabilityNoChange (Radarr)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=false and openWorldHint=false, so the safety profile is partly covered. The description adds real behavioral context beyond that: the preview-then-confirm gating for multi-item edits, and the rationale that root folders are excluded because it would move files. It does not say whether applied changes are reversible or how partial failures are reported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then the required workflow, then the exclusion, then examples. The middle sentence is dense but every clause carries instruction. The trailing example list is mildly decorative but grounds the trigger phrases.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter mutation tool this covers the critical decision path: when a token is needed, what the user must approve, and what is out of scope. Return values need not be described since an output schema exists. Only irreversibility/undo behavior is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description still adds meaning by explaining that confirmation_token is conditionally required depending on items-vs-filter selection and that the payload must match the previewed selection and changes. It does not clarify the filter/change pairing semantics of the 16 parameters beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb, resource, and scope ('Apply a change to many series or movies at once'), which immediately separates it from the read-only siblings and from bulk_edit_preview. The agent can tell what the tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use rules: bulk or filter selections must go through bulk_edit_preview and user agreement first, passing the confirmation_token with the same selection and changes; a single user-named item may be changed directly. It also names a hard exclusion (root folders) and gives worked examples.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bulk_edit_previewA
Read-only

Preview a change to many series or movies at once (monitoring, quality profile, tags, series type, minimum availability): which items match and what would change. Changes nothing. Select items by IDs/titles and/or filters (status, monitored, quality profile, tag, year, complete). Show the preview to the user; if they agree, apply it with bulk_edit_apply and the confirmation_token returned here.

Examples: "Which ended series would be unmonitored?", "Preview putting the 4K profile on all 2024 movies"

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoFilter: has this tag
itemsNoSeries/movie IDs or titles to edit
statusNoFilter by status. Sonarr: continuing, ended, upcoming. Radarr: announced, inCinemas, released
serviceYes'sonarr' (series) or 'radarr' (movies)
year_toNoFilter: year <= this
add_tagsNoChange: tags to add (created if missing)
completeNoFilter: all episodes on disk (series) / has file (movies)
monitoredNoFilter: currently monitored or not
year_fromNoFilter: year >= this
remove_tagsNoChange: tags to remove
set_monitoredNoChange: monitor or unmonitor
quality_profileNoFilter: current quality profile name or ID
set_series_typeNoChange (Sonarr)
set_quality_profileNoChange: new quality profile name or ID
set_minimum_availabilityNoChange (Radarr)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, and the description reinforces this with 'Changes nothing'. It adds behavior beyond annotations by explaining the confirmation_token return and the required follow-up with bulk_edit_apply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and the key 'Changes nothing' safety statement, then gives selection guidance, workflow, and concise examples. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 15-parameter preview tool, the description covers selection modes, change types, the read-only guarantee, the confirmation_token output, and the sibling apply tool. With an output schema also present, nothing important is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 15 parameters in detail. The description groups parameters into selection filters and changes, but does not add per-parameter meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: preview a bulk change to series or movies. It distinguishes itself from the sibling bulk_edit_apply by saying it changes nothing and is only a preview.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to show the preview to the user and then apply with bulk_edit_apply using the returned confirmation_token. This gives both the workflow and the alternative tool, leaving no ambiguity about when to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_pathsA
Read-only

Check the path setup that breaks most imports: root folders (accessible, free space), remote path mappings, download clients, whether Sonarr/Radarr can actually see the download clients' output folders, and the related health warnings.

Examples: "Is my remote path mapping right?", "Why does Sonarr say the path does not exist?"

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuine scope detail beyond the annotations by enumerating the inspected surfaces (folder accessibility, free space, path mappings, download-client output visibility) and the fact that it surfaces related health warnings. No auth or rate-limit notes, but the read-only diagnostic nature makes those low-stakes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core sentence is front-loaded with the tool's purpose and the concrete failure mode ('breaks most imports'), followed by a terse list and useful examples. Every element earns its place; the structure is efficient though the list-plus-examples arrangement runs slightly long.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary. For a one-optional-param, read-only diagnostic, the description covers what is checked and the typical user questions, which is sufficient for correct invocation. The only minor gap is not clarifying its relationship to sibling diagnostics like test_download_clients.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'service' parameter is fully documented in the schema itself ('sonarr' or 'radarr'; omit to query both). The description's phrasing 'whether Sonarr/Radarr can actually see...' only loosely implies the both-services default and adds no syntax or format value beyond the schema. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Check') and a well-scoped resource ('the path setup') and enumerates exactly what it inspects: root folders, remote path mappings, download clients, and cross-service folder visibility. This distinguishes it from neighbors like test_download_clients, diagnose_import, and get_health, which the agent can rule out without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Two concrete example questions ('Is my remote path mapping right?', 'Why does Sonarr say the path does not exist?') give clear situational triggers for when to reach for this tool. It stops short of naming alternatives or stating when NOT to use it (e.g., versus test_download_clients), so it is strong but not fully explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diagnose_importA
Read-only

Diagnose why a download is not being imported: combines the queue status messages, Sonarr/Radarr's own rejections for the files in the download, and path visibility, and returns the probable cause with a concrete fix (remote path mapping, permissions, manual match, not an upgrade, sample, fake release…).

Examples: "Why isn't S03E04 importing?", "Why is everything stuck in importBlocked?"

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceYes'sonarr' (series) or 'radarr' (movies)
queue_idNoQueue item ID (from get_queue); omit to check every stuck item

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds what the read actually produces (probable cause plus a concrete fix, with example causes like remote path mapping, permissions, sample, fake release), which is useful context beyond the annotations. It doesn't discuss cost or latency of aggregating multiple data sources, but for a read-only diagnostic that is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then supporting detail and examples; every sentence earns its place. Slightly verbose in the parenthetical cause list, but it is informative rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering read-only safety, the description only needs to orient usage, which it does. It is complete enough for a 2-parameter diagnostic tool; only the lack of an explicit alternative-tool pointer keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (service enum, optional queue_id) are documented in the schema, so the baseline of 3 applies. The description reinforces the breadth of what is inspected but adds no syntax or semantics beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Diagnose why a download is not being imported') followed by an enumeration of the sources it combines (queue status messages, Sonarr/Radarr rejections, path visibility) and the kind of answer it produces. An agent can distinguish this from manual_import, get_queue, or find_stalled_downloads without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The examples ('Why isn't S03E04 importing?', 'Why is everything stuck in importBlocked?') give concrete usage triggers, and the queue_id description tells the agent when to scope to one item versus check everything. It stops short of naming the sibling to use instead (e.g., find_stalled_downloads for locating stuck items).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_orphan_downloadsA
Read-only

Importable files in the downloads folder that no queue item is tracking (e.g. Download Station removed the task before Sonarr imported it). Groups them by series/movie and says whether each episode/movie already has a file (duplicate) or is missing. Also flags several files for the same episode. If no folder is configured, the download folders are detected from the queue and the import history. The scan can take a while on a NAS.

Examples: "Why wasn't S03E04 of The Capture imported?", "Is anything stuck in /downloads?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
folderNoFolder to scan as seen by Sonarr/Radarr; default ARR_DOWNLOADS_PATH or auto-detected
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, and the description adds meaningful behavior beyond that: grouping by series/movie, duplicate vs missing classification, multiple-file-per-episode flagging, auto-detection of download folders from queue/history when unset, and a real performance caveat ('can take a while on a NAS'). This is exactly the operational context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core definition, then behavior, then caveat, then examples. Slightly long and the folder auto-detection is stated in both description and schema, but every sentence contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the annotations cover the safety profile. The description covers scope, output grouping, and performance, but does not hint at what the agent should do next (e.g., manual_import) once orphans are found, leaving a minor gap for a read-only audit tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents limit, folder, and service with defaults and bounds. The description only restates the folder auto-detection rule that the schema already states; it adds no new parameter semantics. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('importable files in the downloads folder that no queue item is tracking') and explains the causal scenario behind the gap. An agent can distinguish it from find_stalled_downloads (active downloads that stalled) and diagnose_import (diagosing a specific import failure) without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The examples ('Why wasn't S03E04 of The Capture imported?', 'Is anything stuck in /downloads?') give clear real-world invocation contexts. It does not explicitly name the alternative tools to use when the answer is known to be a stalled download or a full import diagnosis, so routing requires inference, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_stalled_downloadsA
Read-only

Find downloads that are stuck: reported as stalled / no connections / no seeds by the download client, with a download client error, or — after min_hours — still at 0 % (no_progress), downloading with no speed/ETA (no_activity), or with no metadata yet (size 0, typical of magnets without peers). Paused or queued items are not reported. Pair with remove_from_queue(blocklist=true) to grab a different release.

Examples: "Which downloads are stalled?", "Remove the stuck torrents and find another release"

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceNo'sonarr' or 'radarr'; omit to query both
min_hoursNoIgnore items added less than N hours ago (no-progress check)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: the exclusion of paused/queued items, the semantics of the min_hours grace period, and the fact that it flags client errors and metadata-less magnets.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The stall conditions are front-loaded and the follow-up pairing plus examples come last. It is dense but each clause enumerates a distinct detection rule; the example list is slightly redundant given the already-explicit criteria.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained. The description covers detection criteria, exclusions, parameter behavior, and the natural next action, leaving no gap for an agent deciding whether and how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning by tying min_hours specifically to the no-progress check ('after min_hours — still at 0%'), clarifying that the threshold does not gate the other stall conditions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource (find downloads) and enumerates the exact conditions that qualify as 'stalled' (client-reported stall, no connections/seeds, error, 0% after min_hours, no activity, no metadata). This is precise enough to distinguish it from siblings like get_queue and find_orphan_downloads.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It specifies when-not (paused or queued items are not reported) and names a concrete follow-up alternative, remove_from_queue(blocklist=true), plus example user utterances. It stops short of explicitly contrasting against get_queue or find_orphan_downloads as the tool to pick instead, so it is clear but not fully routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_suspicious_releasesA
Read-only

Detect fake/malicious releases, with the indexer that provided them: executables (.exe .scr .lnk .bat .cmd .msi .zipx) in the queue, history and blocklist; grabs made more than a day before the episode aired or the movie's digital/physical release (likely="fake" if they never imported, "early_release" if they did, e.g. box sets); and different downloads for different episodes with the exact same size.

Examples: "Is there any .exe in the queue?", "Which indexer is sending me fake releases?"

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoHow far back to look in history
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare read-only and closed-world, so safety is covered. The description adds domain logic beyond annotations: which indexer sent the release, criteria distinguishing 'fake' vs 'early_release' based on import status, and that it scans queue, history, and blocklist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by compact detection criteria and representative examples; no filler. The semicolon list is dense but each clause adds detection detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return details are unnecessary. For a read-only, two-optional-param diagnostic tool, the description covers what is detected, why, and example queries, leaving no critical gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters (days, service) are already documented in the schema. The description adds no further semantics about them (e.g., default day range or dual-service behavior). Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Detect') and resource ('fake/malicious releases'), and enumerates the three detection patterns (executables in queue/history/blocklist, early grabs, same-size different-episode downloads). Distinguishes from sibling inspection tools like get_history or find_orphan_downloads by its fraud-signal focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Supplies example questions it answers, clarifying intended use, but does not explicitly name when to use get_history or get_blocklist instead or when this tool is inappropriate. Context is clear without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_blocklistA
Read-only

Blocklisted releases with the indexer they came from and the reason.

Examples: "What's in the blocklist?", "Which indexer has the most blocklisted releases?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safe-read profile is covered. The description adds the useful fact that each entry is attributed to an indexer with a reason, but says nothing about ordering, whether the list is capped at the schema's default, or how 'most blocklisted releases' would be derived.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core statement is front-loaded in one sentence and the examples are short. The example block is mildly redundant with the first sentence but does convey intended query shapes, so it earns its place without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not the description's burden, and annotations cover the safety profile. With 100% parameter coverage and no required params, an agent has everything needed to call it; only the lack of alternatives guidance keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both 'limit' and 'service' (including the sonarr/radarr enum and the omit-for-both default) are already documented in the schema. The description adds no syntax or scoping detail beyond that, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the resource (blocklisted releases) and what the result carries (source indexer and reason), which is enough to distinguish it from history, queue, and log tools. It never names an explicit verb like 'list' or 'retrieve', but the resource is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two example questions ('What's in the blocklist?', 'Which indexer has the most blocklisted releases?') imply when an agent would reach for this tool. There is no explicit when-not guidance and no pointer to an alternative for adjacent needs such as history or logs, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_calendarA
Read-only

Upcoming (or recent) episode air dates and movie releases.

Examples: "What airs this week?", "What came out in the last 3 days?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
serviceNo'sonarr' or 'radarr'; omit to query both
days_backNo
days_aheadNo
include_unmonitoredNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, covering the safety profile. The description adds the temporal orientation (forward vs. backward looking) but says nothing about which services are queried, ordering, or how unmonitored items are treated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short lines, purpose front-loaded, with examples appended. No wasted words; every sentence contributes. Structure is clean though slightly terse for a five-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and annotations cover the safety profile. The remaining gap is the undocumented window and monitoring parameters, which an agent must infer from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40%, and the undocumented parameters (days_back, days_ahead, include_unmonitored) receive no explanation in the description. The examples loosely imply time-window semantics but never map to the actual parameters, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (episode air dates and movie releases) with a clear temporal scope (upcoming or recent). No sibling tool overlaps this calendar/release-date function, so an agent can route to it unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two examples ('What airs this week?', 'What came out in the last 3 days?') concretely demonstrate when to reach for this tool. It lacks explicit alternatives or exclusions, but the use case is narrow and the examples remove most ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_command_statusA
Read-only

Status of an asynchronous command (search, refresh, manual import) started earlier. For a finished manual import it also checks that each file really reached the library.

Examples: "Has the search for S03E04 finished?", "Did the manual import work?"

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceYes'sonarr' (series) or 'radarr' (movies)
command_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a genuinely useful non-obvious trait: for a finished manual import it also verifies each file actually reached the library, which an agent could not infer from the schema. It stops short of describing polling/eventual-consistency behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose is front-loaded in the first clause, followed by the manual-import special case and two concrete example questions. No filler, though the examples repeat the intent already stated rather than adding new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the annotations carry the safety profile. The description covers purpose, async nature, and the manual-import verification quirk, leaving only minor gaps such as whether a not-yet-finished command requires re-polling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'service' is documented in the schema (enum values sonarr/radarr), but 'command_id' has no schema description. The description supplies context that command_id refers to an earlier-started command, which is helpful, but adds no format, origin, or retrieval guidance beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns the status of an asynchronous command, and enumerates the command kinds it applies to (search, refresh, manual import). The 'started earlier' framing cleanly separates it from the starter siblings like search_episodes, refresh_item, and manual_import, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'started earlier' phrasing plus the example questions ('Has the search for S03E04 finished?') give clear context for when to call it: only after an async command has been issued and you hold a command_id. No when-not-to-use or explicit sibling routing is provided, which keeps it below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cutoff_unmetB
Read-only

Items that have a file but below the quality profile cutoff (upgrade candidates).

Examples: "What could be upgraded to better quality?", "Which movies are below cutoff?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
serviceNo'sonarr' or 'radarr'; omit to query both
monitored_onlyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered structurally. The description adds the useful scoping fact that only items already having a file are returned, but says nothing about the monitored_only default behavior or result volume. Adequate but not rich, given the lower bar set by annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core definition front-loaded and examples following. No wasted prose, though the examples could be consolidated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and annotations cover the safety profile. The remaining gap is sibling disambiguation from get_missing/search_missing, plus the undocumented monitored_only parameter, leaving the definition just adequate for this query tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: limit and service are documented in the schema, but monitored_only (default true) has no description anywhere. The description adds no parameter-level meaning, so it does not compensate for the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource state ('items that have a file but below the quality profile cutoff') and clarifies intent with the parenthetical 'upgrade candidates'. It is not a tautology and an agent can form a clear picture of what is returned, though it never names the sibling it is distinct from (get_missing).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two examples ('What could be upgraded to better quality?', 'Which movies are below cutoff?') give concrete usage context and map directly to user phrasing. There is no explicit exclusion against get_missing/search_missing, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_disk_spaceA
Read-only

Free and total space of the disks/volumes seen by Sonarr/Radarr.

Examples: "How much disk space is left on the NAS?"

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that results cover disks/volumes 'seen by Sonarr/Radarr,' clarifying scope, but says nothing about which services are queried by default, response timing, or whether the call is cheap or expensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the resource and followed by a concrete example. Nothing is wasted, though the example is illustrative rather than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A read-only, zero-required-parameter tool with an output schema, so return-value explanation is unnecessary. The description plus annotations plus schema cover what an agent needs; only minor gaps remain around default query behavior and result scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: the single optional 'service' parameter is documented in-schema as 'sonarr' or 'radarr' with an omit-to-query-both default. The description mentions both services but adds no format or behavior detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Free and total space of the disks/volumes seen by Sonarr/Radarr.' An agent can tell it reports storage capacity rather than system health or queue state. It does not explicitly distinguish itself from near-neighbors like get_system_status or get_health, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The example question ('How much disk space is left on the NAS?') implies when an agent would reach for this tool, but there is no explicit when-to-use, when-not-to-use, or named alternative. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_episodesA
Read-only

Episodes of a series with has_file, monitored, air date and file quality.

Examples: "Is S03E04 of The Capture downloaded?", "Which episodes of season 2 are missing?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
seasonNoSeason number; omit for all seasons
seriesYesSeries ID or title (e.g. 'The Capture')
only_missingNoOnly aired episodes without a file

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered structurally. The description adds useful context by naming the fields returned, but says nothing about ordering, default result size, or how the limit interacts with season filtering; an output schema exists so return format is not its burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what the tool returns and followed by concrete examples; nothing is padded. It could be marginally tighter, but every element carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering the read-only/closed-world profile, a full output schema, and 100% parameter coverage, the description only needs to frame the purpose and it does so with examples. The one real gap is the absence of any routing against overlapping siblings such as search_episodes or get_missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so series, season, limit, and only_missing are all documented in the schema itself. The description adds no syntax, format, or behavioral detail about these parameters beyond what the schema already states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (episodes of a series) and enumerates the notable return fields (has_file, monitored, air date, file quality), so an agent knows what it gets back. It is a noun phrase rather than a clean verb+resource statement and it never distinguishes itself from siblings like get_series, search_episodes, or search_season, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two examples imply usage contexts (checking whether a specific episode is downloaded, listing missing episodes of a season), which is genuine if indirect guidance. There is no explicit when-to-use versus when-not, and no alternative sibling is named even though get_missing and search_episodes overlap heavily.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_healthA
Read-only

Health warnings and errors reported by Sonarr/Radarr (and Prowlarr if configured): unreachable download client, missing root folder, indexers failing… each with the usual fix when it is a known check.

Examples: "Any health warnings?", "Is something wrong with Radarr?"

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real behavioral context beyond that: the kinds of checks surfaced and that known checks come with the usual fix. It doesn't say whether results are cached or how the Prowlarr case is triggered, but the addition over annotations is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and scope, then issue examples and sample queries. No filler sentences, though the trailing example-utterance block is slightly less information-dense than the opening.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover the safety profile. The description supplies the issue taxonomy and fix hint. The one loose end is the Prowlarr mention that the enum does not support, which could mislead an agent about available service values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single optional parameter, so the baseline is 3. The description adds no guidance on the service parameter and in fact mentions Prowlarr, which is not a valid enum value ('sonarr' or 'radarr'), creating mild ambiguity rather than clarification.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: it reports health warnings and errors from Sonarr/Radarr. The concrete enumeration of issue types (unreachable download client, missing root folder, failing indexers) makes the scope unambiguous and separates it from siblings like get_logs or get_system_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The example utterances ('Any health warnings?', 'Is something wrong with Radarr?') imply the intended context, but the description never states when to prefer this over get_logs, get_system_status, or diagnose_import, nor any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_historyA
Read-only

Activity history: grabs (with indexer), imports, failed downloads, deletions.

Examples: "Which indexer did S03E04 of The Capture come from?", "What failed to download in the last 7 days?"

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoOnly the last N days
limitNoMax items to return
movieNoRadarr movie ID or title
seriesNoSonarr series ID or title
serviceNo'sonarr' or 'radarr'; omit to query both
event_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that grabbed events carry indexer attribution, which is genuine behavioral detail, but says nothing about pagination, default limit of 50, or how far back history is retained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose in one line, followed by two short illustrative examples. The event list partially duplicates the event_type enum in the schema, but the whole thing is tight and no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, and it correctly stays focused on query intent. For a zero-required-parameter query tool with six optional filters, it covers the main axes adequately, though it doesn't mention default time window or result caps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents days, limit, movie, series, service, and event_type. The description adds no syntax or format hints (e.g., title vs. ID resolution for movie/series), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Activity history') and enumerates the event categories it returns (grabs with indexer, imports, failed downloads, deletions), which cleanly separates it from get_logs, get_queue, and get_blocklist. It stops short of naming a sibling it is not, so the differentiation is inferable rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two example questions ('Which indexer did S03E04 come from?', 'What failed to download in the last 7 days?') implicitly show when to reach for this tool, which is more than nothing. But there is no explicit when-not guidance and no pointer to alternatives like get_logs for raw events or get_blocklist for blocked items.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_logsA
Read-only

Latest log entries at the given level (default: errors).

Examples: "Show me Sonarr's recent errors", "Why is Radarr failing? check the logs"

ParametersJSON Schema
NameRequiredDescriptionDefault
levelNoerror
limitNo
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the level default, but says nothing about the 20-entry cap, whether results are truncated, or how "latest" is ordered — modest added value beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-sentence purpose followed by two short, concrete examples that double as intent signals for Sonarr/Radarr phrasing. No filler, though the examples could have been spent on the undocumented limit parameter instead.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the service filter is covered by the schema. The remaining gap is the unspecified result cap/pagination behavior, which is minor for a simple read-only log fetch.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% — the `service` parameter is documented in-schema, while `level` and `limit` rely on enums/defaults alone. The description compensates for `level` by stating the default is errors, but adds nothing for `limit` or its 1–200 range, leaving the coverage gap only partly filled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Latest log entries") plus the scoping dimension ("at the given level") and the default. An agent can tell this apart from siblings like get_health or get_system_status, though the description never explicitly contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two example queries ("Show me Sonarr's recent errors", "Why is Radarr failing? check the logs") imply a troubleshooting use case, which is useful context. However, there is no explicit when-to-use/when-not guidance and no named alternative such as get_health or diagnose_import for diagnosing failures.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_missingA
Read-only

Wanted/missing: monitored episodes (Sonarr) or movies (Radarr) without a file.

Examples: "What episodes am I missing?", "Which movies are still not downloaded?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
serviceNo'sonarr' or 'radarr'; omit to query both
monitored_onlyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds the semantic definition ('without a file') but says nothing about scope of results, ordering, or how this differs from a cutoff-unmet list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short lines: the definition is front-loaded and the examples earn their place by mapping natural-language questions to the tool. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and annotations cover safety. The definition is nearly complete, missing only sibling disambiguation against get_cutoff_unmet.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; the description explains the 'monitored' concept and the sonarr/radarr duality, but omits the undocumented monitored_only parameter and the limit cap. Baseline 3 is appropriate since the schema carries most of the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb/resource ('Wanted/missing: monitored episodes (Sonarr) or movies (Radarr) without a file') and adds concrete example queries, so an agent can immediately tell this apart from siblings like get_cutoff_unmet or search_missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The example questions imply when to reach for this tool, but there is no explicit guidance on when NOT to use it or which sibling (e.g. get_cutoff_unmet, search_missing) covers an overlapping need. Usage is suggested, not routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_queueA
Read-only

Download queue with tracked state (downloading, importBlocked, importPending, failed…), status messages explaining why an import is stuck, indexer and download client. Downloads that Sonarr/Radarr did not grab themselves are hidden (as in the web UI) and only counted, unless include_unknown=true. Stalled downloads (no activity, no metadata, no progress for 6 h) are marked with stall_reason and count as problems, even when Sonarr/Radarr report them as healthy.

Examples: "What's stuck in the queue and why?", "Is S03E04 of The Capture still downloading?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
serviceNo'sonarr' or 'radarr'; omit to query both
only_problemsNoOnly items with warnings/errors or blocked imports
include_unknownNoAlso list downloads in the client that Sonarr/Radarr did not grab

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover readOnly/openWorld safety, but the description adds substantive behavior: unknown downloads are hidden and only counted unless include_unknown=true, and stalled downloads are flagged with stall_reason and treated as problems even when the service reports healthy. This is exactly the kind of non-obvious semantics an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and state model, then filtering rules, then examples. Generally tight, though the stall-detection sentence is dense and the parenthetical list runs long; still, every clause carries signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value details are rightly omitted. The description covers filtering semantics, stall/problem classification, and default hiding of unknown downloads, giving an agent enough to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3; the description still adds value by explaining the real consequence of include_unknown (visibility vs. mere counting) and clarifying what counts as a problem for only_problems, going beyond the schema's terse hints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (download queue) with the tracked states it exposes, plus indexer and download client context. An agent immediately understands this differs from plain history or library tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete example questions ('What's stuck in the queue and why?') that signal intended usage, and explains the include_unknown and stall conditions that shape results. It does not, however, name sibling alternatives such as find_stalled_downloads or diagnose_import, leaving some routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_seriesA
Read-only

Details of one series in Sonarr, including per-season statistics (episodes on disk vs aired, monitored seasons, size).

Examples: "How complete is The Capture?", "Which seasons of Severance are monitored?"

ParametersJSON Schema
NameRequiredDescriptionDefault
seriesYesSeries ID or title (e.g. 'The Capture')

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the content shape (disk vs aired counts, monitored seasons, size), but that largely overlaps with the output schema that already exists, and no further behavioral traits (auth, caching, cost) are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, followed by a compact field list and two short illustrative examples. Every sentence earns its place, though the examples slightly pad the text without adding new invocation detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only lookup with an output schema, the description is complete enough: it explains scope and the main content returned. Only minor gaps remain around what happens with an unknown/ambiguous title.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'series' parameter already documents that it accepts an ID or title with an example. The description adds no additional semantics beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Details of one series in Sonarr') and enumerates what is returned (per-season stats, monitored seasons, size). It is clearly distinct from siblings like get_episodes or search_library, though it never explicitly names an alternative to disambiguate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two example questions ('How complete is The Capture?', 'Which seasons of Severance are monitored?') imply when the tool is useful, giving reasonable context. However, there is no explicit when-to-use/when-not guidance or routing to sibling tools such as get_episodes for episode-level data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_system_statusA
Read-only

Version and reachability of every configured service (Sonarr, Radarr, Prowlarr).

Use it first when something fails, to check connectivity and API keys. Examples: "Are Sonarr and Radarr up?", "Which Radarr version am I running?"

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already establish that this is a safe, closed-world read-only operation. The description adds useful behavioral context beyond annotations, specifically that it reports reachability and checks API keys, and that it should be used first during failure diagnosis.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief, front-loaded and well structured: purpose first, usage guidance second, then concrete examples. Every sentence adds useful information with no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, has an output schema, and has annotations covering safety, the description is complete enough for an agent to select and invoke it correctly. Return values need not be explained because an output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to document. With no parameters, the baseline is 4, and the description does not need to compensate for missing schema details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the resource and scope: version and reachability of every configured service, with examples naming Sonarr, Radarr and Prowlarr. It is specific enough to distinguish this from most sibling tools, but it does not explicitly contrast with close siblings such as get_health or test_download_clients.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: use it first when something fails, to check connectivity and API keys. That is strong guidance for a troubleshooting starting point, but it does not name alternatives or conditions where another tool should be used instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grab_releaseA

Send a specific release (from explain_search) to the download client, even if it was rejected. Only works within ~30 minutes of the search; search again if it fails.

Examples: "Grab that 1080p release anyway", "Download the second one from the list"

ParametersJSON Schema
NameRequiredDescriptionDefault
guidYesRelease guid from explain_search
serviceYes'sonarr' (series) or 'radarr' (movies)
indexer_idYesindexer_id from explain_search (Sonarr/Radarr's id, not Prowlarr's)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say this is not read-only, not destructive, and closed-world. The description adds real behavioral context the annotations can't: the ~30-minute staleness window and the fact that it can override a prior rejection. Failure handling ('search again if it fails') is also disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences front-load the action and the constraint, and the trailing user-phrase examples aid intent matching without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and annotations cover the safety profile. The remaining agent-critical facts — where the identifiers come from and how long the grab stays valid — are all stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so guid, service, and indexer_id are already documented in the schema, including the Sonarr/Radarr distinction and the Prowlarr caveat. The description only repeats the explain_search provenance already present in the schema, adding no new parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (send/grab) and resource (release), plus the key behavioral nuance that it works 'even if it was rejected'. It explicitly anchors itself to the sibling tool explain_search as the source of the release, so an agent can distinguish it from search_* and preview_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete operational constraint ('only works within ~30 minutes of the search; search again if it fails') and implies the precondition that a prior explain_search must have run. It doesn't state an explicit when-not case beyond that, but the routing to explain_search is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_reportB
Read-only

Disk usage and cleanup report: total size, biggest series/movies, unmonitored items that still have files, ended series that are complete, size per quality (movies) and free space per root folder. Read-only: it never deletes anything.

Examples: "What takes the most space?", "What could I delete to free space?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, and the description's "Read-only: it never deletes anything" largely restates that safety profile. It does add one useful clarification for a cleanup-oriented tool: the deletion candidates it surfaces are advisory only and nothing is removed. No mention of cost, pagination, or output shape, so it is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The report contents are front-loaded in one dense sentence, followed by the read-only guarantee and two short example queries. Every line earns its place, though the section enumeration is long enough to read as a list rather than a summary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the report's coverage well. The only real hole is the undocumented `limit` parameter, which matters for interpreting "biggest series/movies".

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: the schema documents `service` but `limit` has only a default and min/max bounds with no stated meaning. The description never mentions `limit` or its role (top-N of the biggest series/movies), so the ambiguity around what "biggest" returns is not resolved by prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific deliverable (a disk-usage/cleanup report) and enumerates its sections: total size, biggest series/movies, unmonitored items with files, completed ended series, size per quality, free space per root folder. That enumeration implicitly separates it from raw-infrastructure siblings like get_disk_space, but no sibling is named explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is conveyed only through example questions ("What takes the most space?", "What could I delete to free space?"), which implies the right context without stating when to prefer this over get_disk_space or search_library. There are no exclusions or prerequisites given, so guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manual_importA
Destructive

Import specific files from the downloads folder into the library (all or nothing). Never imports a whole folder blindly: every file must be listed, must be matched to a series/episode or movie, must have no rejections, and no two files may target the same episode or movie. If anything fails, nothing is imported and the problems are returned. When Sonarr/Radarr cannot identify a file ("Unknown Series", "matched by ID, manual import required"), pass matches to assign it by hand; the assignment is re-validated first. Replacing a file that is already in the library needs allow_replace=true. By default it waits for the import and verifies each file (Sonarr/Radarr report the command as successful even when a file fails to move), returning the reason from their log for any file that did not get in.

Examples: "Import /downloads/The.Capture.S03E04.1080p.mkv", "Import that file as episode 4 of season 3 of The Capture"

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoWait (up to 2 min) for the import and check that every file really got in
filesYesExact file paths as seen by Sonarr/Radarr (from find_orphan_downloads)
matchesNoManual identification for files that were not matched (or were matched wrongly)
serviceYes'sonarr' (series) or 'radarr' (movies)
import_modeNoauto (default, like the web UI): copy/hardlink if the download client is still seeding, move otherwise. copy: always copy/hardlink. move: always move (breaks seeding).auto
allow_replaceNoAllow replacing files that are already in the library (they are deleted)
allow_downgradeNoAlso import files Sonarr/Radarr reject as 'not an upgrade' (lower quality or score than the current file). Only when the user explicitly wants the lower version; needs allow_replace.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this destructive, but the description goes well beyond them: it explains the all-or-nothing validation contract (every file must be listed, matched, rejection-free, and uniquely targeted), that failures return the problems, that assignment is re-validated, that Sonarr/Radarr falsely report success so each file is verified against the log, and that import_mode 'move' breaks seeding. This is unusually rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the all-or-nothing contract are front-loaded, then edge cases, then examples. Dense but nearly every sentence carries a distinct operational fact; it is longer than strictly necessary but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary; the description covers the remaining risk surface (atomicity, validation, replacement, verification, seeding impact) for a 7-parameter destructive tool. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema does not: allow_replace is required before a replacement, and allow_downgrade both needs allow_replace and should only be used on explicit user intent. It also clarifies that `wait` verification exists because the upstream command lies about success.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb (Import), resource (specific files from the downloads folder), and destination (the library), plus the atomicity constraint. It is clearly distinguishable from siblings like diagnose_import (read-only diagnosis) and find_orphan_downloads (discovery of candidate paths).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit conditions are given for the non-obvious paths: pass `matches` when Sonarr/Radarr report 'Unknown Series' or 'matched by ID, manual import required'; allow_replace=true when replacing a library file; allow_downgrade only when the user explicitly wants a lower version. It also names find_orphan_downloads as the source of valid paths, and the two examples show the intended invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

parse_releaseA
Read-only

Show how Sonarr/Radarr interpret a release name: matched series/episodes or movie, quality, languages, custom formats and score. If the item already has a file, compares both and estimates whether it would be an upgrade under the quality profile and why.

Examples: "How would Sonarr score 'The.Capture.S03E04.2160p.WEB.DV-GRP'?", "Why is this release 'not an upgrade'?"

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesRelease name, e.g. 'The.Capture.S03E04.1080p.WEB.H264-NTb'
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: it performs an upgrade comparison against the existing file under the quality profile and explains the reasoning, which the annotations cannot convey. Missing only minor details like whether the comparison requires the item to be in the local library.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core behavior and followed by illustrative examples that earn their place by mapping to real user questions. Slightly more text than strictly necessary, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not document return values, and it still previews the fields the agent will get back. For a 2-parameter read-only tool with full schema coverage, nothing an agent needs in order to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (title, service) are already documented in the schema, including the 'omit to query both' semantics for service. The description's example release name adds only marginal value over the schema's own example, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Show how Sonarr/Radarr interpret a release name') and enumerates the exact outputs (matched series/episodes or movie, quality, languages, custom formats and score). This cleanly separates it from write-oriented siblings like grab_release or rename_files, so an agent can select it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two example questions ('How would Sonarr score ...?' and 'Why is this release "not an upgrade"?') give clear scenario context for when to reach for this tool, plus the conditional 'if the item already has a file, compares both' tells the agent a prerequisite. However, no sibling alternatives are named or excluded, so routing guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_renameA
Read-only

Preview which files would be renamed by the current naming settings (current path -> new path). Nothing is changed. Show it to the user; if they agree, apply with rename_files and the confirmation_token returned here.

Examples: "Which files of The Capture don't follow my naming scheme?", "Preview renaming Dune"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
movieNoRadarr movie ID or title
seasonNoOnly this season
seriesNoSonarr series ID or title

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, and the description reinforces 'Nothing is changed' rather than contradicting it. It adds real context beyond the annotations by disclosing that a confirmation_token is returned and is required for the downstream rename_files call, which is workflow-critical. It stops short of describing the preview payload structure in detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and the non-mutating guarantee in the first two sentences, then adds workflow and examples. The examples earn their place for intent matching, though the parenthetical 'current path -> new path' is the only mildly redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the full preview-then-apply loop including the token handoff. Given 4 optional params and a rich output schema, this is close to complete; only the meaning of limit and the no-filter default behavior are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, with movie, series, and season already documented inline; the description adds no parameter-level meaning (e.g., that limit caps the number of previewed files, or that omitting movie/series previews the whole library). Baseline 3 is appropriate since the schema carries most of the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (preview which files would be renamed) plus the scope (current naming settings, current path -> new path), and implicitly distinguishes itself from the sibling rename_files by being the non-mutating step. An agent can tell exactly what it produces without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: show the preview to the user, and if they agree, apply with rename_files using the confirmation_token returned here. Names the alternative tool and the condition that selects it, plus two example user phrasings that map intent to this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refresh_itemA

Refresh metadata and rescan files on disk for a series or movie.

Examples: "Refresh The Capture", "Rescan the files of Dune"

ParametersJSON Schema
NameRequiredDescriptionDefault
movieNoMovie ID or title
seriesNoSeries ID or title

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so the agent knows this mutates internal state without destroying data. The description usefully adds that it touches files on disk and refreshes metadata, but says nothing about side effects such as triggering metadata provider calls, re-downloads, or runtime cost. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with the core action front-loaded and short examples that aid invocation. The examples add practical value rather than filler, though they slightly duplicate the parameter semantics already in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a mutation tool with two optional and mutually ambiguous target parameters, the description never states whether one of movie/series is required or what the tool does when neither is given, leaving a real operational gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both 'movie' and 'series' are already documented as ID-or-title. The description only restates the same targeting via examples and does not clarify how to choose between the two optional parameters or what happens if neither is supplied. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource pair ('Refresh metadata and rescan files on disk') and scopes it to 'a series or movie', so an agent can immediately distinguish it from siblings like rename_files, search_library, or diagnose_import. The two example utterances reinforce the intent without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the examples show natural-language phrasings but there is no explicit statement of when to prefer this over alternatives such as search_library or rename_files, and no prerequisites (e.g., item must already exist in the library) are given. Adequate but leaves the when-to-use decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_from_queueA
Destructive

Remove items from the download queue, optionally deleting them from the download client and blocklisting the release. With blocklist=true and skip_redownload=false, Sonarr/Radarr automatically search for another release. Removing a torrent from the client stops seeding (private trackers may count it as hit-and-run). Many items at once need confirm=true.

Examples: "Remove the .exe releases from the queue, blocklist them and find another version", "Delete queue item 42 from Download Station"

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoRequired when removing many items at once
serviceYes'sonarr' (series) or 'radarr' (movies)
blocklistNoAdd the release to the blocklist
queue_idsYesQueue item IDs (from get_queue)
skip_redownloadNoWith blocklist=true, do NOT search for a replacement automatically
remove_from_clientNoAlso delete the task/files in the download client

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the destructiveHint annotation by disclosing real consequences: removing a torrent from the client stops seeding, private trackers may count it as hit-and-run, and the blocklist/skip_redownload interaction triggers a replacement search. This is exactly the kind of consequence detail an agent needs before a destructive call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then layers the conditional behaviors, then closes with examples. Every sentence carries operational weight with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the annotations already carry the safety profile. The description covers the remaining gaps: side effects, parameter interactions, and the confirm safeguard, so nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds genuine cross-parameter semantics: the blocklist=true + skip_redownload=false combination triggers an automatic replacement search, and confirm=true is required for multi-item removal. This interaction logic is not inferable from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Remove items from the download queue') and immediately enumerates the optional side effects (delete from client, blocklist). The purpose is unambiguous and clearly distinct from the read-oriented sibling get_queue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete context for when the various behaviors apply (blocklist+skip_redownload triggers an auto-search; many items require confirm) and provides two worked examples of user intent. It stops short of explicitly naming alternative tools or stating when NOT to use it, but the operational context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_filesA

Rename the files of a series (optionally one season) or a movie to match the current naming settings. Call preview_rename first, show it to the user and only call this once they agree, passing the preview's confirmation_token (required for more than one file).

Examples: "Yes, rename the files of The Capture season 3", "Apply the rename to Dune"

ParametersJSON Schema
NameRequiredDescriptionDefault
movieNoRadarr movie ID or title
seasonNoOnly this season
seriesNoSonarr series ID or title
confirmation_tokenNoToken from preview_rename; required to rename more than one file

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the description's job is to add context — which it does via the mandatory preview/consent workflow and the token gate. It still doesn't state reversibility or what happens on a partial failure, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The mutation and its prerequisite are front-loaded in the first sentence, followed by the precise gating instruction and short user-utterance examples. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the description fully covers the risky part of the call — the preview/consent/token contract. Nothing an agent needs to invoke this safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description goes slightly further by framing the target as either a series (optionally narrowed to one season) or a movie, implying the movie/series choice is exclusive — a relationship the schema itself never states. The token's conditional requirement is repeated from the schema rather than newly added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (rename) and resource scope (files of a series/season or a movie) plus the target (current naming settings). The sibling preview_rename is implicitly distinguished by the staged workflow, so an agent can tell which one mutates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit sequencing: call preview_rename first, show it to the user, and only invoke this after agreement with the confirmation_token. It also states the token is required for more than one file, so the when/when-not condition is fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_episodesA

Trigger an indexer search for specific episodes (by ID, or series + season + episode numbers). Episodes searched in the last hour are skipped (force=true to override) and searches over many episodes ask for confirmation first, to respect indexer limits.

Examples: "Search S03E04 of The Capture again", "Look for episodes 3 and 4 of season 1 of Severance"

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoAlso search items Sonarr/Radarr searched very recently
seasonNo
seriesNoSeries ID or title, with season + episodes
confirmNoRequired when the search covers many items (shows the count first)
episodesNoEpisode numbers within `season`
episode_idsNoSonarr episode IDs

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly=false, destructive=false, openWorld=false; the description adds substantive traits not present there: the one-hour dedupe window, the force override, and the confirmation gate for large searches to respect indexer limits. That is genuinely useful behavior disclosure, though it says nothing about failure modes or what a confirmation prompt returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then constraints, then two natural-language invocation examples that help an agent map user phrasing to parameters. Slightly redundant between the confirm sentence and the schema's own confirm description, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the skip/force/confirm semantics cover the main surprises for a six-parameter search tool. The only gap is that no sibling alternative is named, which matters in a catalog containing search_season and search_missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the baseline is 3; the description goes beyond it by explaining that inputs come in two mutually exclusive shapes (episode_ids versus series + season + episodes), which the anyOf/nullable schema does not make obvious. It adds little on confirm beyond restating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Trigger an indexer search for specific episodes") and immediately names the two acceptable input shapes (by ID, or series + season + episode numbers). It does not, however, differentiate itself from close siblings such as search_season or search_missing, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives real operating context (searches within the last hour are skipped, force overrides, large searches require confirmation), but that is behavioral rather than routing guidance. It never says when to pick this over search_season, search_missing, or search_library, so usage versus alternatives is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_libraryA
Read-only

Search the local Sonarr/Radarr library (what is already added) by title.

Does not search the internet; use add_series/add_movie for that. Examples: "Do I have The Capture?", "Which Star Wars movies are in my library?"

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to return
queryNoText to look for in titles; empty lists everything
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety and world-boundary are covered structurally. The description reinforces the local-only boundary ('does not search the internet'), which is useful confirmation, but adds nothing about result ordering, truncation, or what the limit cap implies behaviorally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines: purpose, exclusion/alternative, and examples. The scoping constraint and the routing hint are front-loaded, and every sentence carries information an agent needs. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only local search tool with full schema coverage and an output schema present, the description covers purpose, boundary, alternatives, and usage examples. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters (query, service, limit) are already documented in the schema itself. The description's 'by title' adds a marginal hint about what the query matches against, but provides no syntax, matching rules, or service-switching guidance beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search), resource (local Sonarr/Radarr library), and scope qualifier (what is already added) in the first sentence. It also distinguishes itself from the internet-facing siblings by name (add_series/add_movie), so an agent can route without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the negative case ('Does not search the internet') and names the alternative tools to use instead. The two natural-language example queries ('Do I have The Capture?') make the intended invocation context unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_missingA

Trigger a search for every monitored missing episode (Sonarr) and/or movie (Radarr). Sonarr/Radarr normally find new releases through RSS on their own; mass searches hit every indexer once per item and can exceed indexer limits. Over a few items it first returns how many would be searched and needs confirm=true. Prefer narrower searches.

Examples: "Search everything that's missing", "Search all missing episodes of The Capture"

ParametersJSON Schema
NameRequiredDescriptionDefault
seriesNoLimit the Sonarr search to one series
confirmNoRequired when many items would be searched (shows the count first)
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: mass searches hit every indexer once per item and can exceed indexer limits, and over a few items it returns a count first and requires confirm=true. The annotations only give the generic safety profile; the description supplies the actual rate-limit and two-step confirmation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, follows with the important rate-limit/confirmation caveat, closes with the routing hint and concrete examples. No sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, return values needn't be described. The description covers intent, risk, confirmation gating, and alternatives, so an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning to confirm by describing the count-preview gating behavior that the schema only hints at. It does not elaborate on series/service beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (trigger a search) and resource (every monitored missing episode/movie across Sonarr/Radarr). It clearly scopes itself apart from list-style siblings like get_missing and the narrower search_episodes/search_season tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains when this mass-search is appropriate versus relying on normal RSS, and explicitly says 'Prefer narrower searches.' It stops short of naming the specific narrower siblings (search_episodes, search_season), leaving the agent to infer the exact alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_seasonA

Trigger an indexer search for a whole season (season packs or every episode).

Examples: "Search season 2 of The Capture"

ParametersJSON Schema
NameRequiredDescriptionDefault
seasonYes
seriesYesSeries ID or title

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that this kicks off an indexer search rather than a download, but says nothing about whether it is asynchronous, whether results are grabbed automatically, or any rate/queue implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short lines: the action and scope come first, the illustrative example second. No filler and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a tool that initiates a multi-episode search, the description omits what happens after the search (queued grabs, async results) and the valid season value domain, leaving gaps an agent would need to guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'series' is documented as 'Series ID or title' while 'season' has no description. The example implicitly shows season as a numeric value and series as a title, adding modest meaning, but the accepted season formats (e.g. 0 for specials) remain unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Trigger an indexer search for a whole season') and clarifies scope with 'season packs or every episode,' which implicitly separates it from the episode-level sibling. It never names search_episodes explicitly, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The example ('Search season 2 of The Capture') gives a concrete usage context, but there is no explicit when-to-use/when-not guidance and no mention of search_episodes as the alternative for single episodes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_monitoredA
Idempotent

Monitor or unmonitor episodes (by ID, or series + season + episode numbers in one call), a whole season, a whole series, or a movie.

Examples: "Stop monitoring S03E02 of Lioness", "Stop monitoring season 1 of The Capture", "Unmonitor the movie Cats"

ParametersJSON Schema
NameRequiredDescriptionDefault
movieNoRadarr movie ID or title
seasonNoWith `series`: only this season
seriesNoSeries ID or title
episodesNoWith `series` + `season`: only these episode numbers
monitoredYes
episode_idsNoSonarr episode IDs

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that the call can both monitor and unmonitor (a set-to-boolean operation, consistent with idempotency), but says nothing about required permissions, side effects on child items when monitoring a whole series/season, or refresh behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The scope sentence is front-loaded and dense, followed by three short examples that each demonstrate a distinct granularity. No filler or repetition; every line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and annotations cover safety. With 6 parameters at 83% schema coverage plus combination semantics in the description, an agent has enough to call correctly, though the movie-vs-series duality (Radarr/Sonarr) is only implied through parameter names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 83% schema coverage the baseline is 3, but the description adds genuine value by explaining how parameters combine: episode IDs alone, or series+season+episode numbers 'in one call'. This combination logic goes beyond the per-parameter schema notes and helps the agent construct valid calls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb pair (monitor/unmonitor) and enumerates every supported resource granularity: episode by ID, series+season+episode, whole season, whole series, and movie. An agent can distinguish this from siblings like search_episodes or add_series without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is grounded with three concrete example utterances that map natural-language phrasing to parameter combinations, making the intended invocation pattern clear. However, it never states when NOT to use this tool or names an alternative sibling (e.g., bulk_edit_apply) for batch scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test_download_clientsA
Read-only

Test the connection to every download client configured in Sonarr/Radarr and report which ones fail and why.

Examples: "Can Sonarr reach my download client?", "Test the download clients"

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceNo'sonarr' or 'radarr'; omit to query both

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds that it tests *every* configured client and reports failures with reasons, which is useful scope context, but says nothing about latency, timeouts, or remote-call side effects implied by openWorldHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence and the examples are compact and high-signal. Slightly verbose line wrapping and the pair of examples add length without new information, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers scope, behavior, and failure reporting adequately for a read-only diagnostic. Missing only edge-case context such as behavior when no download clients are configured.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'service' parameter is fully documented in the schema (including its default and 'omit to query both' behavior). The description adds no parameter detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Test the connection') and resource ('every download client configured in Sonarr/Radarr') plus the outcome ('report which ones fail and why'). This is clearly distinguishable from sibling diagnostics like get_health or get_system_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The quoted examples ('Can Sonarr reach my download client?', 'Test the download clients') give concrete trigger phrases for when to reach for this tool. It does not, however, name an alternative such as get_health or state when *not* to use it, so it falls short of explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 36 tool updatesv0.2.0
    • First observedadd_series
    • First observedbulk_edit_apply
    • First observedbulk_edit_preview
    • First observedcheck_paths
    • First observeddiagnose_import
    • First observedexplain_search
    • First observedfind_orphan_downloads
    • First observedfind_stalled_downloads
    • First observedfind_suspicious_releases
    • First observedget_blocklist
    • First observedget_calendar
    • First observedget_command_status
    • First observedget_cutoff_unmet
    • First observedget_disk_space
    • First observedget_episodes
    • First observedget_health
    • First observedget_history
    • First observedget_logs
    • First observedget_missing
    • First observedget_queue
    • First observedget_series
    • First observedget_system_status
    • First observedgrab_release
    • First observedlibrary_report
    • First observedmanual_import
    • First observedparse_release
    • First observedpreview_rename
    • First observedrefresh_item
    • First observedremove_from_queue
    • First observedrename_files
    • First observedsearch_episodes
    • First observedsearch_library
    • First observedsearch_missing
    • First observedsearch_season
    • First observedset_monitored
    • First observedtest_download_clients

TDQS

A3.6/5.0

Scored across 36 tools

Disambiguation4/5

Most tools have clearly distinct purposes and detailed descriptions help differentiate (e.g., parse_release vs explain_search vs grab_release). However, several diagnostic tools overlap: find_stalled_downloads vs get_queue's stall detection, check_paths vs get_health, and library_report vs get_disk_space, creating some risk of misselection.

Naming Consistency4/5

Almost all tools use snake_case with consistent verb_noun patterns (get_queue, search_episodes, add_series). Minor deviations like library_report (noun_report) and manual_import (adjective_noun) are readable but slightly break the convention.

Tool Count2/5

36 tools far exceeds the typical 3-15 well-scoped range and surpasses the 25+ threshold for 'too many'. While the domain is broad, many tools could be consolidated or scoped down without losing core functionality.

Completeness3/5

The surface is extensive for diagnostics and many management tasks (queue, history, health, imports, renaming, bulk edits). But notable gaps exist: no add_movie tool despite Radarr support, and no way to delete a series/movie from the library, which can block common lifecycle operations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables interaction with the *arr media management suite (Sonarr, Radarr, Lidarr, Prowlarr, SABnzbd) and TRaSH Guides through MCP tools, allowing media library management, searching, and configuration via natural language.
    70
    6 npm
    1
    MIT
  • A
    license
    C
    quality
    A
    maintenance
    Enables full control of Sonarr from Claude.ai and Claude Code by exposing all 234 v3 API operations as tools for managing media libraries.
    235
    30 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables Claude Desktop to troubleshoot and manage a self-hosted Plex + Servarr media stack, including diagnosing missing media, monitoring download queues and indexer health, and coordinating operational tasks across Radarr, Sonarr, SABnzbd, qBittorrent, Tautulli, TMDb, Prowlarr, and remote hosts.
    17
    MIT