Skip to main content
Glama

search_homologs

Identify homologous proteins by searching local MMseqs2 databases like Swiss-Prot or PDB; returns top hits with identity, E-value, coverage, organism, and a TSV of all hits.

Instructions

Find homologous proteins with MMseqs2 (Swiss-Prot, PDB, ...).

Searches a local database (default Swiss-Prot). Returns the best hits per query with identity, E-value, coverage, organism and description; the full hit table is written to a TSV file.

Runs as a background job; may return a job_id to poll with get_job.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
top_nNoHits per query shown inline.
evalueNoE-value cutoff.
databaseNoInstalled database name (see list_databases).swissprot
max_hitsNoMax hits per query kept in the output file.
sequenceYesProtein (or DNA, searched translated) sequence(s), raw or FASTA.
sensitivityNoMMseqs2 -s: 1 fast … 7.5 most sensitive.
min_coverageNoMinimum alignment coverage of query and target.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the search source (local DB, default Swiss-Prot), the inline return shape (best hits with identity, E-value, coverage, organism, description), the side effect of writing a full hit table to a TSV file, and the async job semantics with polling via get_job. It omits any note on permissions, resource cost, or runtime expectations for a sensitive search.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short, front-loaded lines: purpose first, then return shape, then async behavior. No filler, and the job_id/get_job caveat is placed where an agent will see it before deciding to call.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, yet the description still usefully notes the TSV side output and the job_id polling path. It is essentially complete for invocation; only multi-sequence batching semantics and resource expectations are left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters are already documented (top_n, evalue, max_hits, database, sequence, sensitivity, min_coverage). The description adds only loose framing ('best hits per query', the returned fields) rather than explaining interactions such as top_n vs max_hits or how sensitivity/min_coverage affect results. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Find homologous proteins') plus the underlying engine (MMseqs2) and database scope (Swiss-Prot, PDB, ...). This plainly distinguishes it from siblings like seq_stats, translate_sequence, and find_orfs, which do not perform homology searches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: it searches a *local* database, defaults to Swiss-Prot, and runs as a background job that may return a job_id to poll with get_job. The get_job routing is exactly the alternative-selection guidance an agent needs. It doesn't state when-not to use it or explicitly point at list_databases for valid DB names in the description body.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.