Skip to main content
Glama
musharna

plant-genomics-mcp

by musharna

VEP: Variant Effect

vep_annotate
Read-onlyIdempotent

Annotate a plant variant from an Ensembl region and alternate allele; returns primary molecular consequence, IMPACT, and SIFT scores per overlapping transcript.

Instructions

Predict a variant's molecular consequences with Ensembl VEP (rest.ensembl.org; free, no key). Variant-first (not locus-first): supply an Ensembl region (chr:start-end:strand, e.g. '1:10000-10000:1') and an alternate allele (e.g. 'C'); returns the most-severe consequence plus one row per overlapping transcript (consequence terms, IMPACT, and SIFT when the variant is coding-missense; the polyphen fields stay null, as Ensembl runs PolyPhen for human only). found=false when Ensembl reports no overlapping feature. Works for all 12 organisms. Defaults to arabidopsis_thaliana; pass organism= for other species.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
alleleYesAlternate allele, e.g. 'C' (or 'A/C', an insertion, etc.)
regionYesEnsembl region chr:start-end:strand, e.g. '1:10000-10000:1'
organismNoPlant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxidarabidopsis_thaliana

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
endNo
foundYesTrue if VEP returned an overlapping feature
inputNoVEP echo of the parsed input
startNo
alleleYesAlternate allele, e.g. 'C'
regionYesEnsembl region, e.g. '1:10000-10000:1'
organismYesResolved Ensembl species slug
allele_stringNo
assembly_nameNoAssembly the call is against
seq_region_nameNo
most_severe_consequenceNoMost severe SO term
transcript_consequencesNoPer-transcript {gene_id, transcript_id, consequence_terms, impact, sift_*, …}

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.18.2

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld/non-destructive, and the description adds substantial behavior beyond them: free with no key, return shape (most-severe consequence plus one row per transcript), SIFT only when coding-missense, polyphen fields null because Ensembl runs it for human only, and found=false on no overlapping feature. This is rich, non-obvious operational detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then constraints, then output behavior, then defaults. Dense but every clause carries information; slightly long relative to a three-parameter tool but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values needn't be re-explained, yet the description still summarizes them helpfully, and it covers the organism scope, the null-polyphen caveat, and the not-found signal. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The region and allele format examples largely duplicate the schema strings; the only genuinely additive parameter context is that organism accepts slug/scientific/common name/taxid across 12 organisms and defaults to arabidopsis_thaliana. Useful but marginal over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Predict a variant's molecular consequences with Ensembl VEP') and immediately frames the scope as variant-first rather than locus-first, which is exactly the axis that separates it from the many locus_* siblings. An agent can tell what it produces (consequence terms, IMPACT, SIFT rows) without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: variant-first not locus-first, region+allele required, default organism, all 12 organisms supported. It stops short of naming a specific alternative sibling tool for the locus case, but the positive framing and the contrast clause make the when-to-use condition unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.