Skip to main content
Glama
musharna

plant-genomics-mcp

by musharna

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
PLANT_GENOMICS_MCP_CACHE_TTLNoPer-backend TTL+LRU cache entry lifetime, in seconds (200-only responses).600
PLANT_GENOMICS_MCP_HTTP_HOSTNoHTTP bind address.127.0.0.1
PLANT_GENOMICS_MCP_HTTP_JSONNo0 switches the response shape to streaming SSE events.1
PLANT_GENOMICS_MCP_HTTP_PORTNoHTTP TCP port.8765
PLANT_GENOMICS_MCP_CACHE_SIZENoMax entries per backend before LRU eviction.256
PLANT_GENOMICS_MCP_HTTP_TOKENNoBearer token for HTTP transport; must be ≥32 chars or HTTP server aborts at startup.
PLANT_GENOMICS_MCP_NCBI_EMAILNoNCBI etiquette contact for BLAST queries. Unset → placeholder + per-call warning; NCBI may throttle.
PLANT_GENOMICS_MCP_HTTP_MAX_BODYNoReject POSTs with Content-Length larger than this (in bytes).2097152
PLANT_GENOMICS_MCP_CACHE_DISABLEDNoAny non-empty value makes every cache a no-op.
PLANT_GENOMICS_MCP_HTTP_STATELESSNo0 keeps per-client session state (SSE-style).1
PLANT_GENOMICS_MCP_BLAST_CONCURRENCYNoMax in-flight BLAST searches per process (NCBI per-IP rate limit).2

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
ensembl_plants_lookup_locusA

Fetch metadata for a plant locus identifier from Ensembl Plants. Defaults to arabidopsis_thaliana; pass organism= for other plant species (oryza_sativa, zea_mays, ...). Locus is the TAIR-style identifier (e.g. AT1G01010 for Arabidopsis NAC001).

get_gene_xrefsA

Fetch cross-database references (UniProt, NCBI Gene, TAIR, ArrayExpress, …) for a plant locus from Ensembl Plants. Defaults to arabidopsis_thaliana; pass organism= for other Ensembl Plants species. Returns count + raw xref list + a by_db rollup keyed on Ensembl's dbname (e.g. 'Uniprot_gn', 'EntrezGene') for fast lookup of a single foreign identifier.

get_sequenceA

Fetch a locus's sequence from Ensembl Plants. seq_type is one of genomic / cds / cdna / protein (default protein — the canonical-transcript product). Closes the lookup → fetch → BLAST loop: feed the returned sequence straight to blast_sequence (protein for blastp, cds/cdna for blastn). Defaults to arabidopsis_thaliana; pass organism= for other plant species.

ensembl_region_queryA

List features overlapping a genomic interval via Ensembl Plants /overlap/region. region is the seq-region name (chromosome / contig, e.g. '1'); start and end are 1-based inclusive. feature is one of gene / transcript / cds / exon (default gene). Answers 'what genes are in this QTL interval / assembly window' without a per-locus lookup. Ensembl caps the span — oversized regions error. ensembl_plants_assembly lists an organism's region names and their lengths; a region outside that list, or a start past its length, is an error here. Defaults to arabidopsis_thaliana; pass organism= for other species.

phytozome_lookup_locusA

Fetch a gene record from Phytozome BioMart (phytozome-next.jgi.doe.gov). Defaults to arabidopsis_thaliana; pass organism= for other Phytozome proteomes (slug, scientific/common name, or NCBI taxid — e.g. glycine_max, sorghum_bicolor). Locus is the source-genome gene name (e.g. AT1G01010, Glyma.01G000100). Returns organism_name, gene_name, chromosome, gene_start, gene_end, strand, description.

resolve_locus_to_uniprotA

Resolve a plant locus to its canonical UniProtKB record. Prefers reviewed (Swiss-Prot) entries; falls back to unreviewed (TrEMBL) when no curated record exists (common for non-Arabidopsis plants). organism accepts a canonical slug, scientific/common name, or NCBI taxid (default arabidopsis_thaliana; e.g. oryza_sativa, zea_mays). A gene symbol answers only when it names one locus; a symbol shared by several loci (ARF1) is InvalidArguments listing them. Returns primaryAccession, uniProtkbId, entryType, recommendedName, geneNames, organism, taxonId, sequenceLength, web_url (a link for the client to open; no tool on this server retrieves it). This is the protein-side entry point — pair with InterPro / AlphaFold / Reactome / structural-bio tools.

locus_literatureA

Search Europe PMC for literature mentioning a plant locus. Free, no API key. Returns up to size results (default 10, capped at 25) with title, authors, journal, year, DOI, PMID, open-access status, citation count, abstract, and the article's web_url (a link for the client to open; no tool on this server retrieves it). For non-Arabidopsis species the species common name is appended to the query to disambiguate locus IDs (rice, maize, ...). Pair with resolve_locus_to_uniprot or ensembl_plants_lookup_locus to ground the locus before fanning out to the literature.

locus_go_annotationsA

Fetch Gene Ontology annotations for a plant locus from QuickGO (EBI). Free, no API key. The locus is first resolved to a UniProt accession via the same logic as resolve_locus_to_uniprot, then QuickGO is queried by geneProductId. Returns annotations[] with goId/goName/goAspect/qualifier/evidence + a by_aspect rollup ({molecular_function: [{goId, goName}, ...], biological_process: [...], cellular_component: [...]}) deduped on goId so the high-level term set is one read away; by_aspect_deduped_on names that key in the payload. truncated is true when numberOfHits exceeds returned — raise limit (max 100).

locus_plant_ontologyA

Fetch Plant Ontology (PO) + Trait Ontology (TO) + experimental-condition (PECO) annotations for a plant locus from Planteome (browser.planteome.org, AmiGO2/GOlr; free, no API key). Complements locus_go_annotations: QuickGO serves GO (species-agnostic), Planteome serves the plant-specific ontologies — PO (anatomy + developmental stage), TO (traits). The locus is matched across Planteome's searchable bioentity fields and filtered by the organism's NCBI taxon. Returns annotations[] (term_id / term_name / ontology / aspect / evidence / reference) + a by_ontology rollup ({PO: [{term_id, term_name}, ...], TO: [...], PECO: [...]}) deduped on term_id. Planteome names genes by these locus ids for arabidopsis, rice, wheat and tomato only; other organisms are refused, and a gene Planteome has no record of is not found. Defaults to arabidopsis_thaliana; pass organism= for other species.

go_enrichmentA

GO + KEGG over-representation analysis for a gene LIST via g:Profiler g:GOSt (biit.cs.ut.ee/gprofiler; free, no API key). Unlike locus_go_annotations (one locus → its terms), this answers 'what is my gene SET enriched for?' — the dominant question for a differential-expression or co-expression cluster. loci is the query gene list (e.g. AT-codes for Arabidopsis, RAP-DB IDs for rice). sources defaults to GO:BP/GO:MF/GO:CC + KEGG; user_threshold is the g:SCS-corrected significance cutoff (default 0.05). Optional background sets a custom statistical domain (default: all annotated genes). Returns enriched[] (term_id/name/p_value/intersection_size/…, capped at top_n by p-value) plus unmapped[] — query loci g:Profiler could not recognize, surfaced so a locus-namespace mismatch is visible. Defaults to arabidopsis_thaliana; pass organism= for any of the 12 species.

gramene_homologsA

Fetch orthologs and paralogs for a plant locus from Gramene compara (data.gramene.org v69). Default homology_type='ortholog'; pass 'paralog' for in-species duplicates or 'all' for everything. 'ortholog' includes syntenic_ortholog_* rows; homoeologs (polyploid subgenome copies) come back only under 'all', and excluded_categories counts every category the filter left out. Returns target_locus + homology category (type) + shared gene_tree_id per hit. Rows carry no taxon unless with_organism=true (adds 'organism' per row) or target_organism is given, which filters to one organism before the cap and adds 'organism' per row. Paralogs here are within_species_paralog only: Gramene drops Compara's other_paralog ('ancient paralogues'); ensembl_plants_paralogs lists both. Pair with resolve_locus_to_uniprot for protein-level enrichment and with blast_sequence for sequence similarity discovery.

kegg_pathwaysA

Fetch KEGG pathway memberships for a plant locus from rest.kegg.jp. Returns a list of pathway IDs + names + KEGG category classes the locus participates in. Pairs with locus_go_annotations for the GO-level functional view. Covers: arabidopsis_thaliana, brachypodium_distachyon, glycine_max, hordeum_vulgare, oryza_sativa, populus_trichocarpa, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms). Non-Arabidopsis loci are bridged to the NCBI Entrez Gene ID KEGG indexes (returned as entrez_gene_id). A gene KEGG knows with no pathway memberships is an ok answer with pathways=[]; a gene KEGG has no record of raises NotFoundError. KEGG v118+ is case-sensitive on the locus: pass AGI loci as uppercase.

bar_gene_summaryA

Fetch the BAR (Bio-Analytic Resource, U Toronto) merged ThaleMine + GAIA-aliases summary for an Arabidopsis locus. Returns the TAIR curator summary + Araport11 computational description from /thalemine/gene_information/ together with the NCBI Gene ID and cross-DB aliases (RefSeq, UniProt, TIGR locus-model IDs) from /gaia/aliases/. Arabidopsis only — ThaleMine carries taxon 3702 plus yeast/human for ortholog cross-reference. BAR is keyless and a Global Core Biodata Resource (2023); replaces the v0.9 subscription-gated tair_locus_info stub for the curator-summary use case.

bar_efp_expressionA

Fetch BAR/eFP world-map natural-variation expression for an Arabidopsis locus. Wraps the world-eFP view at /microarray_gene_expression/world_efp/arabidopsis/{locus} — returns expression across ~36 ecotypes (Bay-0, Col-0, Cvi-1, Ler-2, ...) with per-replicate values, control samples, collection lat/lng, and a per-ecotype mean computed client-side. Arabidopsis only. BAR is keyless and a Global Core Biodata Resource (2023).

bar_aiv_interactionsA

Fetch BAR AIV (Arabidopsis Interactions Viewer) interactions for an Arabidopsis or rice locus. Dispatches by organism: Arabidopsis returns curated GRN paper refs from /interactions/get_paper_by_agi/{locus} (PubMed ID, title, image URL, comments, pipe-split tags; the image is a link for the client to fetch, no tool on this server retrieves it); rice returns predicted PPI partners from /interactions/rice/{locus} with Pearson co-expression r (pcc), evidence hits, and quality score. The kind field discriminates the response shape (grn_papers vs ppi_predictions). Rice requires the MSU LOC_Os* locus format — RAP-DB Osg is rejected upstream. Only Arabidopsis and rice are supported by AIV; other organisms raise OrganismNotSupported.

string_interactionsA

Fetch protein-protein interaction partners from STRING-DB (string-db.org). Accepts either a UniProt accession or a locus identifier. A locus goes to STRING's own resolver, except for wheat, whose STRING proteins carry no locus alias: a wheat locus is resolved via UniProt first. Defaults to arabidopsis_thaliana; pass organism= for other plant species (slug, scientific/common name, or NCBI taxid). Returns first-neighbor partners with the combined STRING score plus per-channel sub-scores (experimental, database, textmining, predicted). The argument was called locus_or_accession before #129; that name is still accepted, deprecated.

tair_locus_infoA

Fetch the TAIR curator-vetted Arabidopsis locus summary. Served via BAR/ThaleMine (U Toronto, Global Core Biodata Resource 2023) since TAIR's free per-locus REST API is gated behind a paid Phoenix Bioinformatics subscription. Returns TAIR curator summary + Araport11 computational description + NCBI Gene ID + cross-DB aliases (RefSeq, UniProt, TIGR locus-model IDs). Arabidopsis only. Alias of bar_gene_summary.

plantcyc_locus_infoA

Fetch metabolic annotation for a locus from PlantCyc / the Plant Metabolic Network (pmn.plantcyc.org; free BioCyc web-services API, no key). Walks gene → enzyme → catalyzed reactions → PlantCyc pathways in the organism's PGDB, returning enzymes[] + reactions[] (id/name) + pathways[] (id/name) — the metabolic-pathway view KEGG and GO don't provide. A non-enzymatic gene (e.g. a transcription factor) returns found=false with empty lists, not an error. reaction_count is the true total even when the lists are capped; pathway_count is too, or null (unknown) when the gene catalyzes more reactions than are walked for pathways. 11 organisms have a PGDB (arabidopsis, rice, maize, soybean, grape, poplar, tomato, barley, sorghum, medicago, brachypodium); wheat is not yet mapped. Defaults to arabidopsis_thaliana (AraCyc, the best-curated); pass organism= for other species.

alphafold_structureA

Fetch the AlphaFold DB predicted-structure summary for a locus (alphafold.ebi.ac.uk; free, no key). Resolves the locus → UniProt accession, then returns the predicted model's global mean pLDDT confidence, the per-band pLDDT distribution with each band's pLDDT range (plddt_band_ranges), modelled residue span, latest model version, and mmCIF / PDB / PAE download URLs — links for the client to fetch; no tool on this server retrieves them. A valid protein with no deposited model returns found=false (a normal outcome, not an error); a locus with no UniProt entry raises a typed NotFoundError. Works for all 12 organisms (UniProt-keyed). Complements resolve_locus_to_uniprot (sequence-level) with the structure-level view. Defaults to arabidopsis_thaliana; pass organism= for other species.

experimental_structuresA

Fetch experimentally-solved (X-ray / cryo-EM / NMR) protein structures for a locus from PDBe (www.ebi.ac.uk/pdbe; free, no key). Resolves the locus → UniProt accession, then returns PDBe's best_structures mapping ranked best-first: per entry the PDB id, chain, experimental method, resolution, coverage, and modelled residue span. Most plant proteins have NO deposited structure — that returns found=false (a normal outcome, not an error); a locus with no UniProt entry raises a typed NotFoundError. structure_count is the true total even when the list is capped. Complements alphafold_structure (the predicted view). Works for all 12 organisms (UniProt-keyed). Defaults to arabidopsis_thaliana; pass organism= for other species.

tf_binding_motifsA

Fetch curated transcription-factor DNA binding motifs for a locus from JASPAR (jaspar.elixir.no; free, no key) — the cis-regulatory view. Resolves the locus → UniProt accession + gene symbol, searches JASPAR by symbol scoped to the organism's taxid, then CONFIRMS each candidate by matching the accession against the profile's uniprot_ids. Returns per motif the JASPAR matrix id, TF class/family, assay type (SELEX / ChIP-seq / PBM / DAP-seq), an IUPAC consensus derived from the position-frequency matrix (e.g. CACGTG, the G-box/ABRE core), motif length, PubMed refs, and the SVG sequence_logo and JASPAR web_url — links for the client to fetch; no tool on this server retrieves them. IMPORTANT: JASPAR's name search is fuzzy, so name-similarity hits belonging to a DIFFERENT gene are returned separately in name_only_matches and must NOT be attributed to this locus; only motifs is UniProt-confirmed. found=false means the gene has no curated profile (not a TF, or its family is unprofiled for that species) — a normal outcome, not an error. Use jaspar_motif to retrieve the raw matrix for any matrix_id. Coverage is Arabidopsis-heavy (1236 profiles) and thin elsewhere (maize 131, soybean 91, wheat 58, tomato 51, rice 10; Brachypodium and sorghum have none). Defaults to arabidopsis_thaliana; pass organism= for other species.

jaspar_motifA

Fetch one JASPAR binding profile by matrix id, including its raw position-frequency matrix (PFM: per-base count vectors keyed A/C/G/T) plus TF class/family, assay type, source species, UniProt accessions, PubMed refs, IUPAC consensus, and the sequence-logo URL (a link for the client to fetch; no tool on this server retrieves it). The drill-down companion to tf_binding_motifs, which returns the derived consensus but not the matrix. Accepts a versioned id (MA0570.1) or a bare base id (MA0570, which resolves to the newest version). Unknown ids raise a typed NotFoundError.

experimental_interactionsA

Fetch CURATED EXPERIMENTAL protein/genetic interaction partners for an Arabidopsis locus from ThaleMine (BAR's InterMine instance; free, no key), sourced from BioGRID, IntAct and PSI-MI. Unlike string_interactions (predicted / text-mined, scored) and bar_aiv_interactions (which returns GRN paper references for Arabidopsis, not partner pairs), every partner here carries the actual experimental provenance: detection method (two hybrid, pull down, genetic interference, ...), PSI-MI relationship type, physical vs genetic class, source database, and the PubMed IDs that reported it. ThaleMine emits one row per evidence record, so rows are aggregated to one entry per partner with evidence_count as a crude support signal; partners are ordered by that count. found=false means the gene is real but has no curated interaction on record — a normal outcome; an unknown locus raises a typed NotFoundError. Arabidopsis only (ThaleMine carries genes for taxon 3702; other organisms raise OrganismNotSupported).

locus_gene_rifsA

Fetch curated GeneRIF functional statements for an Arabidopsis locus from ThaleMine (free, no key). A GeneRIF is a one-sentence, manually curated statement of what the gene does, each anchored to the PubMed ID of the publication that demonstrated it — dense, directly citable functional context that GO terms (locus_go_annotations) and raw abstracts (locus_literature) do not provide. Well-studied genes have many: HY5 (AT5G11260) has 114. Upstream order is preserved because ThaleMine supplies no meaningful ranking, so truncated means later statements were cut, not that they were less relevant. found=false means the gene exists but has no GeneRIF; an unknown locus raises a typed NotFoundError. Arabidopsis only.

interpro_domainsA

Fetch the InterPro domain / family architecture for a locus (www.ebi.ac.uk/interpro; free, no key). Resolves the locus → UniProt accession, then returns the protein's InterPro entries — each with accession, name, type (domain / family / homologous_superfamily / …), source_database (Pfam appears here as source_database='pfam', not a separate tool), the integrated InterPro accession, and residue spans — plus a count_by_type rollup. A protein with no annotated domains returns found=true with an empty list; a locus with no UniProt entry raises a typed NotFoundError. domain_count is the true total even when the row list is page-capped. Works for all 12 organisms (UniProt-keyed). Defaults to arabidopsis_thaliana; pass organism= for other species.

locus_variantsA

List natural (germline) variants overlapping a locus's genomic span via Ensembl (rest.ensembl.org; free, no key). Resolves the locus → gene coordinates, then returns EVA/dbSNP-sourced SNPs and indels with id, source, consequence class, alleles, and clinical significance. variant_count is the true overlap total; the variant list is capped for payload size with truncated flagged. Opens the variation axis (distinct from get_sequence / ensembl_region_query). Works for all 12 organisms. Defaults to arabidopsis_thaliana; pass organism= for other species.

vep_annotateA

Predict a variant's molecular consequences with Ensembl VEP (rest.ensembl.org; free, no key). Variant-first (not locus-first): supply an Ensembl region (chr:start-end:strand, e.g. '1:10000-10000:1') and an alternate allele (e.g. 'C'); returns the most-severe consequence plus one row per overlapping transcript (consequence terms, IMPACT, and SIFT when the variant is coding-missense; the polyphen fields stay null, as Ensembl runs PolyPhen for human only). found=false when Ensembl reports no overlapping feature. Works for all 12 organisms. Defaults to arabidopsis_thaliana; pass organism= for other species.

panther_familyA

Fetch the PANTHER protein-family classification for a locus (pantherdb.org; free, no key). Returns the PANTHER family and subfamily (id + name) plus curated GO terms grouped by aspect (molecular_function / biological_process / cellular_component), the PANTHER protein class, and pathways. found=false when PANTHER cannot map the locus; a mapped locus can still have a null family or subfamily, which means PANTHER assigns it none. Complements the sequence-homology tools (gramene_homologs / consensus_homologs) with an evolutionary-family view. Works for all 12 organisms. Defaults to arabidopsis_thaliana; pass organism= for other species.

entry_membersA

List every protein in one organism that carries an InterPro, Pfam or PANTHER entry, with the gene locus each maps to — the reverse of interpro_domains / panther_family (entry -> genes, not gene -> entries). One UniProt query (rest.uniprot.org; free, no key). Each member gives accession, symbol and locus; locus is the id the locus-keyed tools accept, or null when UniProt links the protein to no gene (e.g. an old cDNA submission). reviewed_only=true (default) keeps Swiss-Prot entries: complete for Arabidopsis and rice, empty for most other crops, so pass reviewed_only=false there. total is UniProt's count across all pages; when truncated=true, pass next_cursor back as cursor= for the next page. An entry absent from the organism is ok with total 0. Defaults to arabidopsis_thaliana.

gene_tree_membersA

List the member genes of an Ensembl Compara (plants) gene tree — the gene_tree_id that gramene_homologs returns on every homolog (rest.ensembl.org /genetree; free, no key). Each member gives the locus (the id the locus-keyed tools accept), protein_id, species and taxid, and organism (the canonical slug, or null for a species outside this server's 12). target_organism= keeps one organism's members; omit it for every species. total counts members before limit; truncated=true when limit cut some off. An unknown tree id is a not-found error.

ensembl_plants_paralogsA

List the paralogues Ensembl Compara (plants) records for a plant locus (rest.ensembl.org /homology, type=paralogues; free, no key). Each row gives the paralogue's locus, its type — within_species_paralog, or other_paralog: Ensembl's 'ancient paralogues', inferred across a super tree, so the two genes can sit in different gene trees — the taxonomy_level of the duplication, perc_id/perc_pos and protein_id, closest first. gramene_homologs carries only within_species_paralog, so a gene can have paralogues here and none there. A paralogue list is not a family list: other_paralog can name a gene outside the family (AT2G23390, an acyl-CoA N-acyltransferase-like gene, is an other_paralog of the ARFs); test membership with interpro_domains. An empty list means Compara records no paralogue, not that the gene is single-copy (FLS2, AT5G46330, has none). found=false when Compara keeps no homology record for the gene at all (e.g. a non-coding gene); an id Ensembl does not know is a not-found error. total and counts_by_type count before limit; truncated=true when limit cut some off.

ensembl_plants_assemblyA

Describe an organism's Ensembl assembly (rest.ensembl.org /info/assembly; free, no key): assembly name, GCA accession and date, the karyotype, and every top-level seq-region with its length. The names are the region values ensembl_region_query takes, and a start past a region's length is refused there, so a region walk can be planned before the first call. Karyotype regions come first, in karyotype order, then unplaced scaffolds and contigs, longest first. Names are Ensembl's: tomato's chromosomes are CM001064.4 and so on, not '1'. coord_system labels differ between assemblies (chromosome, scaffold, supercontig, primary_assembly), so in_karyotype, not coord_system, says which regions are chromosomes. total counts every region before limit; truncated=true when limit cut some off (soybean has over 1,100).

upstream_releaseA

Report the release a backend's own release endpoint calls current, for the backends whose answers state none (their tools' upstream_version is always null). It is read by a SEPARATE request at query time, so it is not proof of the release that answered any one call: read it before and after a run, and equal values mean no release changed in between. Never cached. ensembl_plants: the Ensembl Genomes release (e.g. '63'); string: the STRING release ('12.0'); quickgo: the GO annotation load date and the GO ontology date; jaspar: the newest ACTIVE release (several are active at once); kegg: the pathway and genes last-update dates (KEGG has no release number). pdbe, aragwas and europe_pmc publish no data release: release is null and reason says why. Backends whose answers state or pin their release (UniProt, InterPro, PANTHER, AlphaFold, Gramene, ATTED-II, OrthoDB) are not listed: read upstream_version on their answers instead.

orthodb_orthologsA

Resolve a locus to its OrthoDB ortholog group and cross-species member genes (data.orthodb.org; free, no key). Searches at the Viridiplantae level, then returns the group metadata (name, evolutionary rate) and member genes grouped by organism (organism, gene id, description). organism_count is the true cluster total; the member list is capped with truncated flagged. found=false when the locus maps to no ortholog group. Works for all 12 organisms. NOTE: unlike the other locus tools, organism= does NOT scope the search — the group is resolved from the locus id alone at the Viridiplantae level, and organism is only validated and echoed back. Passing a mismatched organism therefore still returns the locus's real group. target_organism= DOES filter: it keeps only that organism's members, before the cap, so a 2,000-member group cannot hide rice or wheat behind 'limit'.

aragwas_associationsA

Fetch AraGWAS genome-wide association study hits for an Arabidopsis locus (aragwas.1001genomes.org; free, no key). Returns each significant SNP association overlapping the gene with its score (-log10 p), minor-allele frequency, the SNP's predicted molecular effect (impact, amino-acid change), and the phenotype/study it came from, including the study's own significance thresholds on the score's scale (study.thresholds), which the over_bonferroni / over_fdr / over_permutation flags are taken against. Rows come strongest first, 25 per answer by default (limit, up to 100); next_cursor resumes at the first row not returned. association_count is the true total even when page-capped. ARABIDOPSIS-ONLY — any other organism raises OrganismNotSupported. Defaults to arabidopsis_thaliana.

arabidopsis_natural_variationA

Fetch 1001 Genomes natural-variation SNP effects for an Arabidopsis locus (tools.1001genomes.org; free, no key) — the variation observed across 1135 resequenced natural accessions. Returns per-SNP effect rows (chromosome, position, accession id, effect, impact, amino-acid change, transcript) plus the gene's genomic span. variant_count is the true row total even when capped. ARABIDOPSIS-ONLY — any other organism raises OrganismNotSupported. Defaults to arabidopsis_thaliana.

batch_ensembl_plants_lookup_locusA

Batch variant of ensembl_plants_lookup_locus. Uses Ensembl's native POST /lookup/id endpoint — one HTTP round-trip for up to 50 loci, materially cheaper than N parallel GETs. Successes in results[] with the same shape as the single-locus tool. Retries 429/5xx via the shared _http helper (Retry-After capped at 60 s). Misses (loci with no record) still land in errors[] with the [NotFoundError] prefix; the whole batch only fails when the retry budget is exhausted.

batch_get_gene_xrefsA

Batch variant of get_gene_xrefs. Fans out per-locus xref lookups over Ensembl Plants in parallel (up to 50 loci). Each results[locus] is the full single-locus shape (count + xrefs[] + by_db rollup).

batch_phytozome_lookup_locusA

Batch variant of phytozome_lookup_locus. Fans out per-locus BioMart queries in parallel (up to 50 loci). Each results[locus] is the full single-locus row (organism_name, gene_name, chromosome, start/end/strand, description).

batch_resolve_locus_to_uniprotA

Batch variant of resolve_locus_to_uniprot. Fans out per-locus UniProtKB searches in parallel (up to 50 loci). Each results[locus] is the full single-locus record (primaryAccession + uniProtkbId + entryType + geneNames + organism + sequenceLength + web_url + …).

batch_locus_literatureA

Batch variant of locus_literature. Fans out per-locus Europe PMC searches in parallel (up to 50 loci). Each results[locus] is the full single-locus payload (query + hitCount + returned + hits[]).

blast_sequenceA

Run a BLAST sequence-similarity search against NCBI BLAST URLAPI. Async Put/Get under the hood — submits the query, polls the RID (honoring NCBI's per-RID 60s floor), and returns the parsed top hits + raw text report excerpt. Programs: blastn / blastp / blastx / tblastn / tblastx. Database defaults to swissprot for protein programs, core_nt for nucleotide. Emits notifications/progress on each poll. Long searches (>10 min) raise [UpstreamUnavailableError] with the RID preserved so the client can re-poll. Set PLANT_GENOMICS_MCP_NCBI_EMAIL to identify the request per NCBI etiquette.

batch_locus_go_annotationsA

Batch variant of locus_go_annotations. Two-stage fanout — each locus is resolved to UniProt and then queried in QuickGO. Per-locus NotFoundError from either stage lands in errors[] with the typed prefix preserved. Capped at 50 loci.

batch_gramene_homologsA

Batch version of gramene_homologs. Up to 50 loci per call; shares the homology_type filter across all loci. Returns the standard batch envelope (count + results dict + errors dict).

batch_kegg_pathwaysA

Batch version of kegg_pathways. Up to 50 loci per call. Covers: arabidopsis_thaliana, brachypodium_distachyon, glycine_max, hordeum_vulgare, oryza_sativa, populus_trichocarpa, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms).

batch_bar_gene_summaryA

Batch variant of bar_gene_summary. Fans out per-locus BAR ThaleMine + GAIA-aliases calls in parallel (up to 50 loci). Each results[locus] is the full single-locus payload (curator summary, computational description, NCBI Gene ID, cross-DB aliases). Arabidopsis only.

batch_bar_aiv_interactionsA

Batch variant of bar_aiv_interactions. Fans out per-locus BAR AIV calls in parallel (up to 50 loci); all loci in a single call share the same organism. Each results[locus] is the full single-locus payload (kind=grn_papers for Arabidopsis with papers list, kind=ppi_predictions for rice with partners list).

batch_string_interactionsB

Batch version of string_interactions. Up to 50 inputs per call. The argument was called loci_or_accessions before #129; that name is still accepted, deprecated.

atted_coexpressionA

Fetch co-expressed gene neighbors from ATTED-II (atted.jp, API v5) for a plant locus. Returns top_n neighbors with target locus + NCBI Entrez gene ID + score (higher = stronger coexpression), in the index the release declares as score_type: 'z' for Ath-u.c4-0, 'LSmr' (logit score) for the other releases. The ATTED-II release (e.g. Ath-u.c4-0 for Arabidopsis, Osa-u.c1-0 for rice) is resolved per-organism. Covers: arabidopsis_thaliana, glycine_max, medicago_truncatula, oryza_sativa, solanum_lycopersicum, vitis_vinifera, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms). A locus that is not in the organism's release raises NotFoundError. Pairs with string_interactions to surface high-confidence functional partners (interactors that are also coexpressed).

batch_atted_coexpressionA

Batch version of atted_coexpression. Up to 50 loci per call. Covers: arabidopsis_thaliana, glycine_max, medicago_truncatula, oryza_sativa, solanum_lycopersicum, vitis_vinifera, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms).

analyze_locus_synthA

Synthesis: one-call equivalent of the analyze_locus prompt. Resolves a locus through Ensembl Plants, then fans out to xrefs, UniProt, Europe PMC, and QuickGO in parallel. Returns a SynthesisEnvelope with per-step status and a reconciled summary flagging cross-source name/accession disagreements.

find_homologs_synthA

Synthesis: one-call equivalent of the find_homologs prompt. Runs BLAST then resolves UniProt-shaped subject accessions via the batch UniProt helper. Returns ranked hits each annotated with their UniProt record (or null if subject_id is not a UniProt accession).

biological_context_synthA

Synthesis: one-call equivalent of the biological_context prompt. Resolves UniProt accession, then fans out to Gramene homologs, KEGG pathways, STRING-DB partners, and ATTED-II coexpression in parallel. Adds a consensus_partners ranking that merges STRING + ATTED scores.

consensus_homologsA

Synthesis: cross-source homology consensus. Resolves UniProt + FASTA sequence, then runs Gramene homology calls and NCBI BLAST in parallel. Dedupes hits by normalized locus token and scores by n_sources * mean_identity — Gramene contributes identity=1.0, BLAST contributes pident/100.

gene_reportA

Synthesis: one-shot 'tell me about this gene' dossier. Resolves a locus through Ensembl Plants + UniProt, then fans out to cross-references, KEGG pathways, STRING interactors, Europe PMC literature, and QuickGO GO terms. Returns a SynthesisEnvelope whose result.markdown is a rendered Markdown gene dossier (the headline output) alongside a structured result.sections mirror; each backend payload appears once, under sections, while steps[] carries status and per-step timing only. The literature section carries no abstracts (abstracts_included: false) and the GO section no withFrom (with_from_included: false); locus_literature and locus_go_annotations return them. result.gene_names labels the Ensembl and UniProt gene names separately when they differ. Any single backend failure degrades that section to an 'Unavailable' note; the rest of the dossier still renders.

batch_locus_callA

Run one locus-keyed tool over up to 50 loci in one call (#131). 'tool' names any tool whose only required argument is 'locus' (interpro_domains, alphafold_structure, panther_family, orthodb_orthologs, gene_report, ...); 'args' holds that tool's other arguments, shared by every locus and checked against its schema once before any call. Returns the standard batch envelope: results keyed by locus, each exactly what the single tool returns, and per-locus errors. The dedicated batch_* tools remain.

Prompts

Interactive templates invoked by user choice

NameDescription
analyze_locusWalk the assistant through a full gene profile for a plant locus: Ensembl annotation, cross-references, UniProt protein record, recent literature, and GO term summary. Chains five tools in a deterministic order.
find_homologsRun a BLAST sequence-similarity search against NCBI and resolve the top hits against Ensembl Plants / UniProt. Chains blast_sequence with the per-hit lookup tools.
biological_contextBuild a biological-context profile for a plant locus by chaining homology (Gramene) → pathways (KEGG, Arabidopsis only) → interactions (STRING) → coexpression (ATTED-II, Arabidopsis only). Cross-references the result lists to surface high-confidence functional partners. For non-Arabidopsis organisms, KEGG + ATTED steps are omitted automatically because those backends only ship Arabidopsis data — the chain still runs Gramene + UniProt + STRING.

Resources

Contextual data attached and managed by the client

NameDescription
Cache statisticsPer-backend TTL+LRU cache stats (hits / misses / size). Sourced from each backend module's process-local _CACHE.
Phytozome organismsMap of canonical slug → Phytozome organism_id, derived from the ORGANISMS registry. Only includes organisms with a non-None phytozome_int. See pgmcp://organisms/coverage for the full coverage matrix across all backends.
Backend statusPer-backend rollup (name, base_url, kind=live, subscription_gated). Lets a client enumerate the live backends without parsing the server docstring.
Organism coverage matrixMarkdown table of all 12 supported plants × 9 ID slots (ncbi_taxid, ensembl, phytozome, string, europe_pmc, kegg, atted, gprofiler, plantcyc). Missing slots render as em-dash. Lets a client introspect coverage in one read instead of probing resolve_organism per organism.

TDQS

A3.7/5.0

Scored across 56 tools

Disambiguation3/5

The descriptions are unusually thorough and often explicitly contrast sibling tools (e.g., string_interactions vs experimental_interactions, alphafold_structure vs experimental_structures). However, the set contains several genuinely overlapping clusters: the tair_locus_info/bar_gene_summary alias, multiple synthesis dossiers (gene_report, analyze_locus_synth, biological_context_synth, consensus_homologs), many homology/orthology tools, and the generic batch_locus_call alongside dedicated batch_* variants. An agent could still misselect among these, especially when choosing between a one-shot dossier and its component tools.

Naming Consistency4/5

Most tools use snake_case with recognizable patterns: batch_* for batch variants, resource_action or locus_attribute for single-source tools, and _synth for prompt-equivalent syntheses. Minor deviations exist, such as verb-first names (get_sequence, resolve_locus_to_uniprot, find_homologs_synth) mixed with noun-first names (locus_go_annotations, interpro_domains, blast_sequence), but the convention remains readable and largely predictable. This is not chaotic, just slightly mixed.

Tool Count2/5

56 tools is far beyond the typical 3-15 range and is inflated by many dedicated batch variants plus synthesis wrappers, even though a generic batch_locus_call already exists. While the server integrates many plant-genomics data sources, the surface feels heavy and contains redundancy that could be consolidated. This is too many tools for an agent to navigate comfortably, despite the breadth of the domain.

Completeness4/5

The surface covers a wide range of plant-genomics needs: locus resolution, sequences, BLAST, GO and plant ontology, pathways, homology, interactions, coexpression, expression, variation, GWAS, structure, domains, literature, and TF motifs. Minor gaps remain, such as no arbitrary genomic-region sequence retrieval, no multiple sequence alignment or phylogenetic tree tool, and several tools restricted to Arabidopsis or a 12-organism list. Still, the read-only research surface is robust overall.

Maintenance

ActivityActive
ResponsivenessResponsive