Skip to main content
Glama
musharna

plant-genomics-mcp

by musharna

BLAST: Sequence Search (NCBI)

blast_sequence
Read-only

Submit a nucleotide or protein sequence to NCBI BLAST to find similar sequences; returns parsed top hits and raw report text.

Instructions

Run a BLAST sequence-similarity search against NCBI BLAST URLAPI. Async Put/Get under the hood — submits the query, polls the RID (honoring NCBI's per-RID 60s floor), and returns the parsed top hits + raw text report excerpt. Programs: blastn / blastp / blastx / tblastn / tblastx. Database defaults to swissprot for protein programs, core_nt for nucleotide. Emits notifications/progress on each poll. Long searches (>10 min) raise [UpstreamUnavailableError] with the RID preserved so the client can re-poll. Set PLANT_GENOMICS_MCP_NCBI_EMAIL to identify the request per NCBI etiquette.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
expectNoE-value threshold (default 10).
programNoBLAST program — default blastp.blastp
databaseNoNCBI BLAST database slug (e.g. swissprot, core_nt, refseq_protein). Defaults to swissprot for protein programs and core_nt for nucleotide programs.
max_waitNoMax seconds to wait for the search to finish before raising UpstreamUnavailableError with the RID preserved (default 600).
sequenceYesRaw or FASTA-formatted query sequence.
megablastNoEnable megablast (blastn only). Default false.
hitlist_sizeNoMax hits to return (default 10).
poll_intervalNoSeconds between polls. Clamped up to NCBI's per-RID 60s floor.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
ridYesNCBI BLAST request ID — re-usable via fetch_result()
hitsYesTop alignments, sorted by BLAST default order
totalNoAlways null: the upstream returns the top hits asked for and states no total; null means unknown, never zero (#123)
statusYesAlways "READY" when this object is returned
programYesblastn | blastp | blastx | tblastn | tblastx
databaseYesNCBI BLAST database, e.g. swissprot, core_nt
hitCountYesNumber of rows parsed from the alignment summary
returnedYesRows in this payload (#123)
truncatedNoAlways null: without a stated total, truncation is unknown (#123)
elapsed_secondsYesWall-clock from submit to READY
raw_report_excerptYesFirst 50 KB of the FORMAT_TYPE=Text report
raw_report_truncatedYesTrue if the upstream report exceeded the cap

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv1.29.0
    • changedInput schema / properties / max_wait / description
      Previous value: -"Max seconds to wait for the search to finish before raising NotFoundError with the RID preserved (default 600)."New value: +"Max seconds to wait for the search to finish before raising UpstreamUnavailableError with the RID preserved (default 600)."
  2. Changed4 schema fields changedv1.22.0
    • addedOutput schema / properties / returned
      Added value: +{
      +  "description": "Rows in this payload (#123)",
      +  "title": "Returned",
      +  "type": "integer"
      +}
    • addedOutput schema / properties / total
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "integer"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Always null: the upstream returns the top hits asked for and states no total; null means unknown, never zero (#123)",
      +  "title": "Total"
      +}
    • addedOutput schema / properties / truncated
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "boolean"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Always null: without a stated total, truncation is unknown (#123)",
      +  "title": "Truncated"
      +}
    • changedOutput schema / required
      Previous value: -[
      -  "rid",
      -  "program",
      -  "database",
      -  "status",
      -  "hitCount",
      -  "hits",
      -  "raw_report_excerpt",
      -  "raw_report_truncated",
      -  "elapsed_seconds"
      -]New value: +[
      +  "returned",
      +  "rid",
      +  "program",
      +  "database",
      +  "status",
      +  "hitCount",
      +  "hits",
      +  "raw_report_excerpt",
      +  "raw_report_truncated",
      +  "elapsed_seconds"
      +]
  3. First observedv1.8.0

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations give only the safety profile (readOnly, openWorld, non-destructive). The description adds substantial behavioral detail beyond them: async Put/Get polling, the NCBI per-RID 60s floor, progress notifications per poll, the >10 min failure mode raising UpstreamUnavailableError with the RID preserved for re-polling, and the required NCBI etiquette env var. This is exactly the operational context an agent needs and justifies a top score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded, opening with the core operation before moving to async mechanics, program/database defaults, and error behavior. Every sentence carries information, though the async/error/env-var block is a touch long. Efficient overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present it need not describe return values, and it doesn't over-explain them. For a long-running, polling-based external call it covers the full behavioral surface: submission, polling cadence, progress events, timeout/failure semantics with RID preservation, and NCBI identification. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters including the database default mapping and the poll floor. The description largely restates the program list and database defaults already present in the schema, adding little semantic value beyond it. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Run a BLAST sequence-similarity search against NCBI BLAST URLAPI.' It even enumerates supported programs, so the agent knows exactly what operation is performed. It stops short of differentiating itself from homology-oriented siblings (find_homologs_synth, consensus_homologs), which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the agent can infer this is the tool for raw/fasta sequence input as opposed to a locus identifier. There is no explicit 'use this when you have a sequence, use X when you have a locus' routing, and no stated exclusions. Adequate but leaves selection vs siblings to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.