Skip to main content
Glama

Metallum MCP

An MCP server that lets Claude (or any MCP client) look things up in Encyclopaedia Metallum (metal-archives.com): bands, lineups over time, discographies, releases, tracklists, lyrics, artists, labels, reviews, similar bands and catalog browsing.

It is built on pymetal, wrapped with a rate-limited client and a few extra fetchers for things pymetal doesn't cover.

Tools (22)

Search

Tool

What it does

search_bands

By name and/or advanced filters: genre, country, formation year range, lyrical themes, location, label. Pages with offset; returns total and next_offset

search_albums

By title, band, year range, release type, genre, label, country. Pages like search_bands

search_songs

By title, band, release, or words in the lyrics

Bands

Tool

What it does

get_band

Profile plus lineup grouped by current / past / live / last-known / guest, with each member's roles and years

get_band_bio

Full biography (the band page only shows an excerpt)

get_discography

Releases with type, year and review stats, optionally filtered by type, year range, or reviewed releases only

band_review_stats

For up to 25 bands at once: release count, review count and review-weighted average score, sorted best first. Filters by type, year range and minimum reviews

get_similar_bands

User-voted similar artists, ranked by votes

get_band_links

Official site, Bandcamp, Spotify, socials, merch

get_band_reviews

Review list (score, reviewer, date, URL)

random_band

A random band, optionally from a genre bucket

Releases

Tool

What it does

get_album

Details, tracklist (per-disc, bonus/instrumental flags, lyrics_id), lineup split into band / guest / staff

get_other_versions

Reissues, remasters and regional editions

get_lyrics

Lyrics by lyrics_id

get_review

Full text of one review

People and labels

Tool

What it does

get_artist

Real name, born/died, cause of death, origin, plus full biography and trivia

get_label

Address, styles, founding date, sub-labels, parent label

Browse and discovery

Tool

What it does

browse_bands

By country code, genre bucket or first letter

browse_reviews

Reviews posted in a given month

get_upcoming_releases

Upcoming releases, soonest first

get_rip_artists

The R.I.P. list

list_countries

Country codes for filters

Typical flow: search_bands → get_band / get_discography → get_album → get_lyrics.

For "which bands are rated highest" questions: page through search_bands, then pass batches of ids to band_review_stats. Averages come from Metal Archives' rounded per-release averages, so treat scores within about half a point as ties.

Related MCP server: steam-mcp

Setup

Requires uv. It installs Python 3.12 for the project; your system Python is not used.

git clone https://github.com/vincejyr/metallum-mcp.git
cd metallum-mcp
uv sync                        # create .venv and install pinned deps
uv run python test/smoke.py    # end-to-end test against the live site (~75s)

Claude Code

claude mcp add --scope user metallum -- "$(pwd)/.venv/bin/metallum-mcp"

Claude Desktop

Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS), using the absolute path to your clone:

{
  "mcpServers": {
    "metallum": {
      "command": "/absolute/path/to/metallum-mcp/.venv/bin/metallum-mcp"
    }
  }
}

How it uses pymetal

  • Pinned to commit ce5d75a (v1.2.0a1, reviewed 2026-10-01). PyPI's pymetal is an older release published from a different repo, so install from git.

  • No anti-bot layer. The antibot extra (unblock_requests: Cloudflare challenge solving, Wayback Machine fallback) is deliberately not installed. Plain requests work fine.

  • Polite client (src/metallum_mcp/client.py) replaces pymetal's HTTP behaviour:

    • It doesn't impersonate Chrome's TLS fingerprint, and sends a fixed, honest User-Agent instead of a random one on every request.

    • It enforces the site's Crawl-delay: 3 across all requests and threads.

    • It caches responses in memory for 30 minutes, and discography pages on disk for 7 days (SQLite at ~/.cache/metallum-mcp/), so review-stat scans survive restarts and don't re-crawl.

  • Browse and review tools fetch one page only (paginate=False). pymetal's default is to walk the entire catalogue.

Gaps in pymetal that this server fills (extras.py)

Gap

Fix

Band.comment is always None (wrong HTML selector)

get_band_bio reads the site's "read more" endpoint

Artist.biography is always None (bio is loaded separately)

get_artist fetches the bio and trivia from the "read more" endpoints

No way to get a review's text

get_review parses the review page

With a country filter, search_bands puts the band's location in country

search_bands moves it to location and fills country from the filter

These are worth reporting upstream.

Settings

Variable

Default

Purpose

MA_MIN_INTERVAL_MS

3000

Minimum gap between requests

MA_CACHE_TTL_MS

1800000

In-memory cache lifetime

MA_DISK_CACHE_TTL_DAYS

7

Disk cache lifetime for discographies (0 disables it)

MA_CACHE_DIR

~/.cache/metallum-mcp

Disk cache location

MA_USER_AGENT

Mozilla/5.0 (compatible; metallum-mcp/1.0; personal use)

User-Agent header

Notes

  • Each uncached request takes about 3 seconds. get_artist with its bio makes 3 requests, and a genre-filtered random_band can take up to about 30 seconds. band_review_stats makes one request per uncached band, so a full batch of 25 takes about 75 seconds the first time.

  • Meant for personal, interactive lookups, not bulk scraping.

  • Parsing depends on the site's HTML. If a tool starts returning empty fields, run the smoke test to see which one broke.

  • Uses MCP Python SDK 2.x (MCPServer, formerly FastMCP).

Evals

evals/ holds an end-to-end eval: 16 real research questions (facts, multi-step lookups, rating rankings, name collisions, edge cases, an ambiguous question), run through headless Claude Code with only this server attached. Graders check the model's final ANSWER: line. One ambiguity case is judged by Sonnet 5.5. Each case also records tool calls, whether band_review_stats was used, tokens, cost and latency. Cases and expected answers are in evals/cases.md.

Site data is replayed from recorded fixtures (MA_FIXTURES + MA_FIXTURE_MODE=record|replay in the client), so scores don't drift as Metal Archives changes and runs don't crawl the site. The fixtures aren't committed. Record your own first (about 4 minutes, live site). This writes to a scratch flow, because the committed baseline's results would make the runner skip every case:

EVAL_FIXTURE_MODE=record node evals/run-eval.mjs --flow .claude/hillclimb/metallum-record --variant baseline --model claude-opus-5-5 --concurrency 1 --approve-harness

After that, runs replay offline (EVAL_FIXTURE_MODE=replay, the default), e.g. --flow .claude/hillclimb/metallum-research --variant v1 to compare a change against the committed baseline. A page that wasn't recorded fails its case as fixture_miss rather than being fetched. Re-run in record mode to fill it in. --approve-harness records a fingerprint of the runner and cases; the runner refuses to run if they change until someone approves again.

Baseline (2026-10-02, Opus 5.5): 16/16 correct, ~$0.76 per full run, median 2 tool calls and ~8 s per case.

Layout

src/metallum_mcp/server.py   MCP server and tool definitions
src/metallum_mcp/client.py   rate-limited, cached HTTP client for pymetal
src/metallum_mcp/extras.py   bios, trivia, review text
test/smoke.py                      end-to-end test through a real MCP client
evals/                       eval runner, cases and grading (see Evals)

Credits and disclaimer

  • Data comes from Encyclopaedia Metallum, maintained by its volunteer community. This project is not affiliated with or endorsed by Metal Archives. Please respect the site and its rate limits.

  • Built on pymetal (Apache-2.0).

License

MIT. See LICENSE.

Available Tools

21 tools
browse_bandsBrowse bandsB
Read-only

List bands from the catalog by country, coarse genre, or first letter (first limit rows).

ParametersJSON Schema
NameRequiredDescriptionDefault
byYes
limitNoMax results to return
valueYesCountry code ('NO'), genre slug ('black'), or letter ('A', 'NBR' for numbers, '~' for other)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, covering the safety profile. The description adds one useful behavioral fact, the '(first limit rows)' truncation, but says nothing about ordering, pagination, or what a band record contains. Given the annotation coverage, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words, and the parenthetical about limit adds real information rather than padding. It could be slightly tighter, but nothing is superfluous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter list tool with no output schema, the description is minimally adequate: it conveys what gets returned and the truncation rule. It omits any detail about result ordering, records returned, or interplay with search_bands, leaving clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is a moderate 67%: limit and value are documented in the schema, while 'by' has only an enum. The description restates the three filter axes and the 'coarse' genre nuance, but adds little syntax or format detail beyond what the enum and value description already provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (List) and resource (bands from the catalog) and enumerates the three filter axes (country, coarse genre, first letter). It is clearly distinguishable from get_band/random_band, though it does not explicitly distinguish itself from the sibling search_bands.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The enumeration-based phrasing ('by country, coarse genre, or first letter') implies a browsing use case as opposed to a keyword lookup, but there is no explicit when-to-use statement or reference to the search_bands alternative. Guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browse_reviewsBrowse reviewsA
Read-only

Reviews posted in a given month (defaults to the current month). Metadata only; use get_review for the text.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNo
limitNoMax results to return
monthNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds genuinely useful behavioral context: results are metadata only (no review text) and the month window defaults to the current month.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the scope/default comes first and the alternative-tool routing follows. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully signals that only metadata is returned, and the alternative for text is named. Minor gaps remain around year semantics and how limit interacts with pagination, but nothing critical for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (limit is documented; year and month are not). The description partially compensates by explaining the month default, but gives no guidance on the year parameter's format or how month/year interact, so the gap is not fully closed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (browse) and resource (reviews) with a temporal scope ('posted in a given month') and the default behavior. Distinguishable from get_review, which it explicitly names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to get_review when the review text is wanted, which is the key decision point among the many siblings. It doesn't cover other alternatives (e.g., get_band_reviews) or when this tool is inappropriate beyond the text distinction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_albumGet releaseA
Read-only

Release details (type, date, label, catalog no., format, review stats, notes), the tracklist (with per-track lyrics_id), and the lineup split into band / guest / staff.

ParametersJSON Schema
NameRequiredDescriptionDefault
album_idYesMetal Archives release id (from search_albums or get_discography)

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered, but with no output schema the description usefully discloses the shape of results: release metadata, tracklist with linkable lyrics_id, and a three-way lineup split. It stops short of noting pagination, error cases, or whether any of the nested data can be null, but the return-structure disclosure is genuinely additive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence front-loads the most important content (release details) before the tracklist and lineup. No filler, though the parenthetical field list is somewhat packed and could be trimmed without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only lookup with no output schema, the description covers the essential return surface (metadata fields, tracklist, lineup categories) that an agent needs to decide the call is worth making. It omits nothing critical, though depth of nested fields remains unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: the album_id field is fully described in the schema (Metal Archives release id, sourced from search_albums or get_discography). The description adds nothing about the parameter, so the baseline 3 for schema-driven params is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (a release/album) and enumerates the return payload: details, tracklist with lyrics_id, and lineup split into band/guest/staff. It reads as a content manifest rather than a verb+resource statement, and it never names siblings like get_lyrics, get_review, or get_other_versions to draw a boundary, so an agent still infers differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative guidance in the description. Nothing tells the agent when to reach for get_album versus search_albums, get_discography, or get_other_versions, leaving routing entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_artistGet artistA
Read-only

Artist profile (real name, born/died, cause of death, origin) plus, by default, the full biography and trivia (two extra requests).

ParametersJSON Schema
NameRequiredDescriptionDefault
artist_idYesMetal Archives artist id (from a lineup)
include_bioNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint, so the description usefully adds cost/latency context: the bio and trivia are 'two extra requests' issued by default. It also discloses the field set for a tool with no output schema. It stops short of describing the response shape in detail or rate-limiting behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the field list and the cost caveat both earn their place. Slightly dense parenthesis usage, but nothing wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries return-value burden and does so by naming the profile fields plus the extra bio/trivia payload. What is missing is what happens when include_bio is false (bare profile only) and any pagination or error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: artist_id is documented as a Metal Archives id from a lineup, while include_bio is bare. The description compensates by explaining that bio/trivia are fetched by default and cost two extra requests, which is exactly the semantics the agent needs to decide on that flag.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (artist profile) and enumerates the returned fields (real name, born/died, cause of death, origin), which distinguishes it from siblings like get_band and get_band_bio. It never names a sibling explicitly, so differentiation requires inference, holding it below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, nor any mention of the alternative tools (get_band, get_band_bio) an agent might pick instead. The 'by default' clause implies the bio fetches automatically but does not tell the agent when to turn it off.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_bandGet bandA
Read-only

Band profile (country, location, status, formation year, genres, lyrical themes, label) plus lineup grouped by status (current, past, live, last_known, guest_session), with each member's roles and years. Use get_band_bio for the biography.

ParametersJSON Schema
NameRequiredDescriptionDefault
band_idYesMetal Archives band id (from search_bands)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds valuable structural detail about what the return contains (lineup grouped by status categories), but says nothing about auth requirements, rate limits, pagination, or error/not-found behavior. With annotations carrying the safety burden, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste: the return-content summary is front-loaded, and the sibling routing closes it. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully carries the return shape (profile fields and status-grouped lineup). It is nearly complete for a single-param detail tool, though it omits behavior on missing/invalid ids.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single band_id parameter is documented there (including 'from search_bands' and exclusiveMinimum), so the schema does the heavy lifting. The description adds no syntax or format detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (get band) and enumerates the returned content: profile fields (country, location, status, formation year, genres, lyrical themes, label) plus lineup grouped by status with roles and years. It also explicitly names get_band_bio as the sibling it is not, so an agent can distinguish it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes a specific sub-use to an alternative: 'Use get_band_bio for the biography.' This is clear when-to-use guidance. It doesn't enumerate other nearby siblings (get_discography, get_band_links, get_band_reviews), though those are functionally distinct, so it stops just short of full coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_band_bioGet band biographyC
Read-only

Full band biography / history text.

ParametersJSON Schema
NameRequiredDescriptionDefault
band_idYesMetal Archives band id (from search_bands)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds only that the text is 'full', with no note on behavior when a bio is absent, whether the text is external/scraped, or any rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the resource immediately front-loaded. It is arguably under-specified rather than over-long, but there is no wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read tool with annotations covering safety and no output schema, this is minimally adequate. It stops short of telling the agent what the return text looks like or how to handle a band with no biography.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single band_id parameter is fully documented in the schema, including its origin (from search_bands). The description adds no syntax or format meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (band biography/history) and states the scope is the full text. It is distinguishable from get_band (metadata) and get_discography, though it does not explicitly contrast itself with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as get_band for summary data. The agent must infer that this fetches the long-form bio rather than any other band attribute.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_band_reviewsGet band reviewsA
Read-only

Review list for a band's releases: release, score %, reviewer, date and review_url. Metadata only — pass review_url to get_review for the full text.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return
band_idYesMetal Archives band id (from search_bands)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds meaningful behavioral context beyond annotations by specifying that it returns metadata only and that review_url must be passed to get_review for full text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is two compact sentences with no wasted words. The return fields and the metadata-only constraint are front-loaded, followed immediately by the routing instruction to get_review.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with annotations and full schema coverage, the description is complete enough. It compensates for the lack of an output schema by listing the returned fields and explains how to obtain the full review text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema. The description adds no additional meaning for band_id or limit, which is acceptable given the complete schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: it returns a review list for a band's releases. It also lists the exact fields returned, which clearly distinguishes it from get_review, the full-text detail tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent to use get_review when the full review text is needed, which is the key alternative for this tool. However, it does not address when to use it versus browse_reviews or any exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_discographyGet discographyB
Read-only

A band's releases with id, type, release year and review stats. Optionally filter by type.

ParametersJSON Schema
NameRequiredDescriptionDefault
band_idYesMetal Archives band id (from search_bands)
release_typesNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and openWorldHint=true, so safety and external-data behavior are covered. The description adds value by listing the returned fields and the optional filter, but does not disclose return format, pagination, or rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loaded with the resource and return fields, followed by the only optional behavior. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only retrieval tool with annotations covering safety and no output schema, the description is nearly complete: it identifies the resource, the returned fields, and the optional filter. Minor gaps remain around pagination or output ordering, but these are not critical for basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: band_id is documented in the schema, while release_types has no description. The description adds only 'optionally filter by type', which marginally clarifies the release_types parameter but does not explain the allowed enum values or the null default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource (a band's discography) and enumerates the returned fields, making the tool clearly distinct from siblings like get_album or get_band. It is clear but does not explicitly name alternatives, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage guidance is that the type filter is optional. There is no explicit guidance on when to use this tool versus siblings like get_album, search_albums, or get_band, nor are prerequisites beyond the schema's band_id note stated in the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_labelGet labelC
Read-only

Record label profile: country, status, address, styles, founding date, sub-labels, parent.

ParametersJSON Schema
NameRequiredDescriptionDefault
label_idYesLabel id (from a band or release)

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, covering the safety profile. The description's added value is the enumeration of returned fields (country, status, address, styles, founding date, sub-labels, parent), which partly substitutes for the missing output schema, but it says nothing about error behavior or what an invalid label_id yields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is extremely short and wastes no words, but it is a dangling fragment rather than a front-loaded sentence, so the brevity comes at the cost of stating the action. The field list is dense but not scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only single-parameter getter with no output schema, the field enumeration gives the agent a reasonable picture of the return payload. It still omits any statement of what the tool does and when to call it, which leaves the definition thin for its role among many sibling retrieval tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the schema documents the single label_id parameter ('Label id (from a band or release)') and its exclusiveMinimum constraint. The description adds no further detail about the identifier, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (record label) and enumerates the profile fields returned, which lets an agent tell it apart from get_artist, get_band, and get_album by resource. However, it is a noun-phrase fragment with no verb ('Get label profile...'), so the actual action is only implied by the tool name rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of the label_id source being required, and no reference to any alternative sibling tool. The agent must infer from the name alone that this is the retrieval call for label entities.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_lyricsGet lyricsC
Read-only

Lyrics for a song.

ParametersJSON Schema
NameRequiredDescriptionDefault
lyrics_idYeslyrics_id from get_album tracks or search_songs

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds nothing further — no note on availability, licensing/region limits, or what is returned when lyrics are unavailable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment with no waste, but that brevity reflects under-specification rather than disciplined conciseness. There is nothing to front-load because nothing meaningful is said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the burden of conveying return behavior, yet it says only 'Lyrics for a song.' For a lookup tool with no output schema, this leaves the agent without any sense of the response shape or failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole parameter's origin (from get_album tracks or search_songs) is fully documented in the schema. Per the baseline rule, high coverage with a silent description still lands at 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Lyrics for a song' essentially restates the tool name and title without adding a specific verb+resource framing beyond what the identifier already conveys. It does not distinguish this tool from siblings like get_album or search_songs beyond the implicit resource name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives. The only routing hint lives in the schema param description ('from get_album tracks or search_songs'), not in the tool description itself, so the description provides no usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_other_versionsGet other versionsC
Read-only

Other versions of a release: reissues, remasters, regional and format editions.

ParametersJSON Schema
NameRequiredDescriptionDefault
album_idYesMetal Archives release id (from search_albums or get_discography)

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true, so the safety and scope profile is already covered. The description adds no behavioral context such as whether results are paginated, whether an empty list is possible for releases without variants, or what each entry contains — it only lists content categories, which is closer to purpose than behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment with no wasted words, but it is under-specified rather than concise: there is no front-loaded verb and no supporting structure, so brevity comes at the cost of clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should ideally hint at what is returned; the category list partially does this. But for a single-param lookup tool with no usage guidance and no return-shape detail, it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (album_id) and schema description coverage is 100%, with the schema itself pointing to search_albums/get_discography as sources of the id. The description adds no parameter meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource ('other versions of a release') and enumerates what those versions are (reissues, remasters, regional and format editions), which helps an agent distinguish it from get_album. However, it is a noun-phrase fragment with no verb, so the action (retrieve/list) must be inferred, and it never mentions that it returns versions for a specific album.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no exclusions. The sibling set includes get_album and get_discography, and an agent would benefit from knowing to call this only after identifying a base release, but the description says nothing about that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_reviewGet reviewA
Read-only

Full text of a single review: title (with score), author, date and body.

ParametersJSON Schema
NameRequiredDescriptionDefault
review_urlYesreview_url from get_band_reviews or browse_reviews

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. With no output schema, the description adds real value by disclosing the shape of the returned data (title with score, author, date, body), though it says nothing about failure modes or whether a review can be missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste; the returned-content clause earns its place because there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only getter with no output schema, the description adequately covers both the resource and the return shape. It is close to complete, missing only edge-case behavior such as what happens for an invalid or deleted review_url.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single required parameter, so the schema already carries the semantics. The description adds nothing about the parameter beyond what the schema and its description state; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('full text of a single review') and enumerates the returned fields (title, score, author, date, body). It is clearly differentiated from list-oriented siblings like browse_reviews and get_band_reviews by the singular 'single review' framing, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the tool fetches one review, so an agent infers it should be used after obtaining a review_url. The routing hint lives in the schema parameter description ('review_url from get_band_reviews or browse_reviews') rather than in the tool description itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_rip_artistsGet deceased artistsA
Read-only

Metal Archives' R.I.P. list: deceased artists with band, date and cause of death.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds useful context by naming the record fields returned, but says nothing about pagination behavior, ordering, or the limit cap beyond what the schema already states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence identifies the source list and the fields returned, with zero filler. Nothing is wasted and the key information comes first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-required-param list tool with no output schema and a 100%-covered input schema, the description supplies enough to call it correctly. It could go slightly further on default result size or whether the list is filterable, but the essentials are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter at 100% schema description coverage, the schema already documents `limit` (default 50, min 1, max 200). The description adds no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Metal Archives' R.I.P. list: deceased artists') and enumerates the returned data (band, date, cause of death). This clearly distinguishes it from the band/album/review siblings, though it never explicitly names an alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is strongly implied by the specific purpose — this is the only R.I.P./deceased-artist listing among the siblings — but there is no explicit when-to-use statement, no prerequisite note, and no routing to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_similar_bandsGet similar bandsB
Read-only

User-voted similar artists, ordered by match score (number of votes).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return
band_idYesMetal Archives band id (from search_bands)

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds genuinely useful behavior (results are user-voted and sorted by vote count, i.e. a popularity proxy rather than similarity algorithm), but says nothing about pagination, limit interaction, or what happens for a band with no votes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single clipped sentence with zero filler, and the most decision-relevant fact (ordering by vote count) is front-loaded. It is efficient, though the verbless fragment form leaves no room for structural cues such as required-input notes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only lookup with full schema coverage and no output schema, the description conveys what is returned and how it is ordered. The main omission is pagination/limit behavior, which is minor for a tool capped at 200 results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so band_id's origin (from search_bands) and limit's bounds/default are already documented in the schema. The description adds no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (similar artists) and its ordering semantics (match score / vote count), which clearly distinguishes it from generic siblings like get_band or get_discography. It stops short of explicitly tying the output to the input band_id, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives such as get_artist or search_bands, nor any note that a band_id must first be obtained. Usage context must be inferred entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_upcoming_releasesGet upcoming releasesC
Read-only

Upcoming releases, soonest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety and external-data profile is covered elsewhere. The description's one addition is the deterministic 'soonest first' ordering, which is genuinely useful, but it omits everything about result shape, pagination beyond the limit param, and data freshness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single five-word fragment with zero waste and the ordering constraint front-loaded. It is underspecified rather than bloated, but as pure structure it is efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should at least sketch what a 'release' contains and whose releases are returned; neither is present. Annotations cover the safety profile, so the gap is moderate rather than severe.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single optional parameter with 100% schema description coverage ('Max results to return', default 50, max 200), so the schema carries the semantics. The description adds nothing about the limit and doesn't need to; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Identifies the resource ('upcoming releases') and the sort order, but the verb is elided and the scope is ambiguous — it never says whose upcoming releases (global feed vs. a specific band/artist). It is distinguishable from the browse_*/search_* siblings but not clearly from get_discography, which also returns a set of releases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this rather than get_discography, get_album, or search_albums, and no prerequisites or exclusions stated. The only implied usage is the sort order, which the agent could infer from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_countriesList countriesA
Read-only

Country codes Metal Archives uses (for country filters and browse_bands).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety and openness profile is covered. The description adds only that the payload is Metal Archives' own country code vocabulary, which is genuinely useful for filtering, but says nothing about format, size, or completeness of the list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short line with zero filler, and the resource is front-loaded ahead of the usage clause. It is a fragment rather than a complete sentence, which slightly weakens it but costs no tokens.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but 'Country codes' plus the stated consumers of those codes tells an agent what it will get and why. It could go one step further and confirm the return shape (e.g., code list only vs. code+name pairs), which is the only real gap for a no-param lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline of 4 applies; there is nothing for the description to disambiguate. The mention of usage contexts (country filters, browse_bands) gives the only semantic anchor relevant to a no-input tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact resource ('Country codes Metal Archives uses'), so an agent knows it returns the canonical code list rather than country data or band results. It is a telegraphic fragment with no verb, but the title plus resource name make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(for country filters and browse_bands)' implies this is an input-lookup for country filter parameters, which is useful routing information. However, it never states explicitly when to call it (e.g., 'call before filtering by country to get valid values') and names no alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

random_bandGet random bandA
Read-only

A random band from the archive, optionally from a coarse genre bucket. Genre-filtered picks re-roll until they match, so they can take up to ~30 seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
genreNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint, openWorldHint), so the bar is lower, yet the description adds a genuinely useful trait not in structured data: genre-filtered picks re-roll until they match and can take up to ~30 seconds. That latency/retry disclosure materially affects how an agent should call and wait on this tool. It stops short of describing the returned band's shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, with the core purpose front-loaded and the caveat/cost second. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-param read tool with readOnlyHint covering safety, the description supplies purpose, optionality, and the key latency caveat. With no output schema, the only minor gap is not indicating what a returned band object contains, but that is a small omission for this tool class.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it partially does by framing 'genre' as a 'coarse genre bucket' and explaining the re-roll consequence of filtering. However, it adds no syntax or semantic detail about the 23 allowed enum values or how they behave at boundaries (e.g., near-miss genres failing), leaving the enum itself to carry the meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('a random band from the archive') with an explicit scope modifier ('optionally from a coarse genre bucket'). This clearly distinguishes it from the deterministic siblings get_band, browse_bands, and search_bands, which an agent can tell apart without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the random/serendipity nature of the tool, and the genre bucket is presented as optional, but the description never says when to prefer this over browse_bands or search_bands, nor when a genre filter is appropriate. No exclusions or explicit alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_albumsSearch releasesC
Read-only

Search releases by title, band name, release year range, type, genre, label or country.

ParametersJSON Schema
NameRequiredDescriptionDefault
bandNo
genreNo
labelNo
limitNoMax results to return
titleNo
countryNoISO country code of the band
year_toNo
year_fromNo
release_typesNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds nothing behavioral beyond that: no default result ordering, no behavior when multiple filters are combined, no note that limit caps at 200, and no indication of pagination or result shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler; the action and the searchable dimensions come first. It is terse to the point of omitting needed context, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, zero-required tool with no output schema and thin annotations, the description should explain how filters combine, default ordering, and what the limit applies to. It covers the filter surface at a minimum-viable level but leaves these gaps open.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 22% across 9 params, and the description partially compensates by naming most facets (title, band, year range, type, genre, label, country) and clarifying that year is a range. But it adds no matching semantics (partial vs exact, case sensitivity) and ignores limit entirely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Search) and resource (releases) and enumerates the facets that can be matched. However, it never distinguishes itself from siblings such as search_bands, search_songs, or get_discography, so an agent must infer which search surface applies to a given query.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The facet list implies what can be filtered, but there is no statement of when to use this over search_bands/get_discography, no note on prerequisite or combining filters, and no exclusions. Usage remains entirely inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_bandsSearch bandsA
Read-only

Search bands by name and/or advanced filters (genre text, country, formation year range, lyrical themes, location, label). At least one filter is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
genreNo
labelNo
limitNoMax results to return
themesNoLyrical themes, e.g. 'occultism'
countryNoISO country code, e.g. 'SE', 'NO', 'US' (see list_countries)
year_toNo
locationNoCity/region text, e.g. 'Gothenburg'
year_fromNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds the non-obvious constraint that at least one filter must be supplied, but says nothing about result ordering, truncation, or what the 200-result limit implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler: the filter inventory comes first and the required-filter constraint is placed last where it reads as a precondition. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter search tool with no output schema and no required parameters, the description covers the filter surface and the critical at-least-one-filter rule. Return shape and pagination semantics are left unstated, which is a minor gap given search conventions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 44% (name, genre, label, year_from, year_to have no schema descriptions), so the description carries extra burden. It partially compensates by labeling genre as free text, tying year_from/year_to to a 'formation year range', and describing themes/location/label, but it adds no syntax, matching behavior, or format details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (bands) and enumerates the filter dimensions (name, genre, country, formation years, themes, location, label), which cleanly separates it from siblings like search_albums, search_songs and browse_bands. It stops short of explicitly contrasting itself with those alternatives, so it falls just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one hard usage rule — 'At least one filter is required' — which is genuinely useful. However it never says when to prefer this over browse_bands or get_band, nor whether an empty result is expected for unknown names, so the routing guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_songsSearch songsA
Read-only

Search songs by title, band, release or lyrics text. Hits include lyrics_id for get_lyrics.

ParametersJSON Schema
NameRequiredDescriptionDefault
bandNo
limitNoMax results to return
titleNo
lyricsNoWords that appear in the lyrics
release_titleNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint and openWorldHint already cover the safety profile, so the bar is lower. The description adds one useful behavioral detail (results carry a lyrics_id chaining to get_lyrics) but says nothing about pagination, result caps, or matching semantics, which the annotations do not cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core purpose front-loaded and the chaining hint second. Nothing redundant or padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description gives the one return detail an agent needs to chain onward (lyrics_id). It omits limit defaults and pagination behavior, but for a read-only multi-field search that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 40% (only 'lyrics' and 'limit' are documented), so the description must compensate. It names title, band, release, and lyrics, covering four of the five inputs and mapping them to the search behavior; only matching syntax (exact vs. substring) is left unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Search songs') plus the searchable dimensions (title, band, release, lyrics). Distinct from siblings like search_albums/search_bands by resource, though it doesn't explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (search by any of these fields) and gives one handoff hint — 'Hits include lyrics_id for get_lyrics' — but never states when to prefer this over search_albums or how to combine filters. Context is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 21 tool updatesv1.0.0
    • First observedbrowse_bands
    • First observedbrowse_reviews
    • First observedget_album
    • First observedget_artist
    • First observedget_band
    • First observedget_band_bio
    • First observedget_band_links
    • First observedget_band_reviews
    • First observedget_discography
    • First observedget_label
    • First observedget_lyrics
    • First observedget_other_versions
    • First observedget_review
    • First observedget_rip_artists
    • First observedget_similar_bands
    • First observedget_upcoming_releases
    • First observedlist_countries
    • First observedrandom_band
    • First observedsearch_albums
    • First observedsearch_bands
    • First observedsearch_songs

TDQS

B3.3/5.0

Scored across 21 tools

Disambiguation4/5

Most tools have clearly distinct purposes, helped by descriptions that cross-reference each other (e.g. get_band vs get_band_bio, get_band_reviews vs browse_reviews). Mild overlap exists between browse_bands and search_bands, and the many band-centric getters could briefly confuse, but each targets a different facet.

Naming Consistency4/5

Nearly all tools follow a consistent snake_case verb_noun pattern (get_, search_, browse_, list_). Only random_band deviates by omitting a verb, a minor inconsistency in an otherwise predictable scheme.

Tool Count4/5

21 tools is on the high side for the typical 3-15 range, but the breadth of the Metallum domain (artists, bands, albums, songs, lyrics, reviews, labels, upcoming, RIP) justifies most of them. Each tool appears to cover a distinct retrieval need, so the count is slightly heavy but reasonable.

Completeness4/5

The read-only surface is comprehensive: artists, bands, albums, songs, lyrics, reviews, labels, countries, upcoming releases, and RIP lists are all covered. Minor gaps include no dedicated search_artists or search_labels, and get_artist does not explicitly return associated bands, but these are workable omissions.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Exposes Steam Web API tools as MCP resources for Claude Code, Claude Desktop, and Gemini CLI, enabling profile lookups, game searches, achievement tracking, and more.
    11
    20 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables extracting metadata from Beatport track URLs, including preview audio, cover art, and track info, via an MCP server integrated with Claude Desktop.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Exposes Metacritic game, movie, TV, and music review data through MCP tools and resources, enabling LLM hosts to search and retrieve reviews with optional filters.
    1
    -