plant-genomics-mcp
Summary: This MCP server gives an LLM 50+ read-only tools to look up plant-genomics data for a locus (default Arabidopsis, 12+ species supported) across 23 free public backends, in single-locus, batch (up to 50 loci), or cross-source synthesis form.
Gene identity & metadata: Ensembl Plants and Phytozome locus records, cross-database xrefs, nucleotide/protein sequences, genomic region feature queries, and assembly/region listings.
Protein & structure: Resolve a locus to UniProtKB; InterPro domain architecture; PANTHER family/subfamily; AlphaFold predicted models and PDBe experimental structures.
Function & pathways: QuickGO GO annotations, Planteome PO/TO terms, KEGG pathways, PlantCyc/PMN metabolic pathways, and g:Profiler GO/KEGG enrichment over a gene list.
Homology & families: Gramene orthologs/paralogs, Ensembl Compara paralogs, gene-tree members, OrthoDB ortholog groups, InterPro/Pfam/PANTHER entry members, and NCBI BLAST sequence search.
Interactions & coexpression: STRING scores, ATTED-II coexpression, BAR AIV predictions, and ThaleMine curated experimental interactions.
Literature & curation: Europe PMC paper search, BAR/ThaleMine curator summaries (via
bar_gene_summary/tair_locus_info), and ThaleMine GeneRIF functional statements.Variation & diversity: Ensembl VEP consequence prediction, EVA/dbSNP variants over a gene, AraGWAS associations, and 1001 Genomes natural variation (Arabidopsis only).
Regulation: JASPAR TF binding motifs per locus and raw motif matrices by matrix ID.
Expression: BAR world-eFP natural-variation expression profiles across Arabidopsis ecotypes.
Batch & synthesis:
batch_*andbatch_locus_callfan out any locus tool over ≤50 loci;gene_reportrenders a Markdown gene dossier;analyze_locus_synth,find_homologs_synth,biological_context_synth, andconsensus_homologscompose multiple backends with per-step status.Provenance:
upstream_releasereports the current release a backend's own endpoint declares; most tools also exposeupstream_versionandtruncated/totalpaging fields.
🌱 plant-genomics-mcp
56 tools for plant-genomics locus lookup over the Model Context Protocol — 29 single-locus + 1 motif lookup + 1 region query + 1 assembly listing + 1 release report + 1 variant annotator + 1 gene-set enrichment + 1 BLAST search + 2 member lists (family entry, gene tree) + 13 parallel-batch + 5 cross-source synthesis variants. 23 free, public sources: Ensembl Plants, Phytozome BioMart, UniProtKB, Europe PMC, QuickGO, Planteome, PlantCyc/PMN, g:Profiler, AlphaFold DB, PDBe, InterPro, JASPAR, PANTHER, OrthoDB, AraGWAS, 1001 Genomes, NCBI BLAST, Gramene, KEGG, STRING-DB, ATTED-II, ThaleMine, and BAR (Bio-Analytic Resource for Plant Biology).
📦 Install
# Zero-install — uv fetches and runs it on demand
claude mcp add plant-genomics --scope local -- uvx plant-genomics-mcp# pipx — installs the CLI onto your PATH
pipx install plant-genomics-mcp
claude mcp add plant-genomics --scope local -- plant-genomics-mcp
# GHCR Docker image
docker pull ghcr.io/musharna/plant-genomics-mcp:latest
claude mcp add plant-genomics --scope local -- \
docker run --rm -i ghcr.io/musharna/plant-genomics-mcp:latest
# From source
git clone https://github.com/musharna/plant-genomics-mcp.git
cd plant-genomics-mcp
python -m venv .venv && .venv/bin/pip install -e .
claude mcp add plant-genomics --scope local -- "$(pwd)/.venv/bin/plant-genomics-mcp"Related MCP server: hgnc-link
💬 Try it
Once connected, ask Claude a plain-language question — you don't have to name any tool or remember the chain:
"Tell me everything about the Arabidopsis gene AT1G01010 — its function, GO terms, KEGG pathways, protein-interaction partners, and recent papers."
The server supplies the tools (here: Ensembl Plants, UniProt, QuickGO, KEGG,
STRING-DB and Europe PMC lookups); which ones get called, in what order and in
how many turns is up to the client. In the recording at the top of this page
(a narrower prompt), Claude Code picked the calls itself and returned one
combined answer. Swap in any locus and pass organism= for cross-species —
e.g. rice Os01g0100100 (oryza_sativa) — and each tool maps the organism to
that backend's own identifier.
A worked 64-call run over 114 genes in three organisms, with the 51 gaps it logged (39 since closed), is in examples/arf_family/PAGE.md.
🛠️ Tools
56 tools across 23 backends — Ensembl Plants, Phytozome BioMart,
UniProtKB, Europe PMC, QuickGO, Planteome, PlantCyc/PMN, g:Profiler,
AlphaFold DB, PDBe, InterPro, JASPAR, PANTHER, OrthoDB, AraGWAS, 1001 Genomes, NCBI BLAST,
Gramene, KEGG, STRING-DB, ATTED-II, ThaleMine, BAR.
29 single-locus + 1 motif lookup + 1 region query + 1 assembly listing + 1 release report + 1 variant annotator + 1 gene-set
enrichment + 1 BLAST search + 2 member lists + 13 parallel-batch + 5 cross-source synthesis. Most take a
TAIR-style locus (e.g. AT1G01010) plus
optional organism= (slug / scientific name / common name / NCBI taxid
— 12-plant curated coverage matrix at the pgmcp://organisms/coverage
MCP resource). All publish JSON outputSchema, EDAM ontology tags, and
behaviour annotations — every tool is readOnlyHint + openWorldHint, so
hosts can surface them without a destructive-action confirmation prompt.
# | Category | Tool | What it does |
1 | Gene metadata (live) |
| Fetches gene record from Ensembl Plants REST (any plant species). |
2 | Cross-references (live) |
| Fetches cross-DB references (UniProt, NCBI Gene, TAIR, GO, …) from Ensembl. |
3 | Gene metadata (live) |
| Fetches gene record from Phytozome BioMart (any Phytozome proteome). |
4 | Protein (live) |
| Resolves a locus to its UniProtKB record (Swiss-Prot preferred, TrEMBL OK). |
5 | Literature (live) |
| Searches Europe PMC for papers mentioning the locus (free, no API key). |
6 | GO annotations (live) |
| Fetches QuickGO GO annotations (locus → UniProt → QuickGO). |
7 | Sequence search (live) |
| NCBI BLAST URLAPI — async Put/Get polling with progress notifications. |
8 | Homology (live) |
| Fetches Gramene v69 homology entries (ortholog / paralog) with gene_tree_id. |
9 | Pathways (live) |
| Fetches KEGG pathway memberships. 7 organisms: Arabidopsis ( |
10 | Interactions (live) |
| Fetches STRING-DB first-neighbor interaction partners with per-channel score. |
11 | Coexpression (live) |
| Fetches ATTED-II top-N coexpression neighbors with a score: z-score for Arabidopsis, logit score (LSmr) for the other releases. |
12 | Curator summary (live) |
| Fetches BAR ThaleMine + GAIA-aliases curator summary for an Arabidopsis locus. |
13 | Expression (live) |
| Fetches BAR eFP-Browser expression profile (mean ± SD per tissue) for a locus. |
14 | Interactions (live) |
| Fetches BAR AIV interaction partners (Arabidopsis + rice) with confidence + papers. |
15 | Curator summary (live) |
| Silent upgrade — alias of |
16 | Metabolism (live) |
| Walks gene → enzyme → reactions → PlantCyc/PMN pathways (free BioCyc web-services API). The metabolic-pathway view KEGG/GO lack; found=false for non-enzymatic genes. 11 species have a PGDB. |
17 | Sequence (live) |
| Fetches a locus's sequence (genomic / cds / cdna / protein) from Ensembl |
18 | Region query (live) |
| Lists gene/transcript/cds/exon features overlapping a genomic interval (chr:start-end) via Ensembl |
19 | Enrichment (live) |
| GO + KEGG over-representation for a gene list via g:Profiler g:GOSt — "what is my DE / co-expression set enriched for?" Reports unmapped loci; optional custom background. All 12 organisms. |
20 | Plant ontology (live) |
| Plant Ontology (anatomy / dev-stage) + Trait Ontology annotations for a locus via Planteome (Solr) — the plant-specific ontologies GO doesn't cover. by_ontology rollup; taxon-filtered. Strong for 6 species. |
21 | Structure (live) |
| AlphaFold DB predicted 3D model for a locus (locus → UniProt → model): global mean pLDDT, per-band confidence, modelled span, and mmCIF / PDB / PAE URLs. found=false when no model is deposited. All 12 organisms. |
22 | Structure (live) |
| PDBe experimentally-solved (X-ray / cryo-EM / NMR) structures for a locus (locus → UniProt): best-first PDB id, chain, method, resolution, coverage, residue span. found=false when none deposited (common for plants). All 12 organisms. |
23 | Domains (live) |
| InterPro domain / family architecture (locus → UniProt): each entry's accession, name, type, source_database (Pfam included), integrated InterPro id, and residue spans, plus a count_by_type rollup. All 12 organisms. |
24 | TF motifs (live) |
| JASPAR curated TF DNA-binding profiles for a locus (locus → UniProt → symbol search, then UniProt-confirmed): matrix id, TF class/family, assay type (SELEX / ChIP-seq / PBM / DAP-seq), IUPAC consensus, PubMed refs, logo URL. Fuzzy name hits for other genes are quarantined in |
25 | TF motifs (live) |
| One JASPAR profile by matrix id (e.g. |
26 | Interactions (live) |
| ThaleMine CURATED EXPERIMENTAL interaction partners (BioGRID / IntAct / PSI-MI) for an Arabidopsis locus — per partner: detection method (two hybrid, pull down, ...), PSI-MI relationship type, physical vs genetic, source DB, PubMed IDs, and an evidence count. The experimental counterpart to |
27 | Function (live) |
| ThaleMine curated GeneRIF statements — one-sentence, manually curated descriptions of what the gene does, each tied to a PubMed ID (HY5 has 114). Citable functional context that GO terms and raw abstracts don't provide. Arabidopsis only. |
28 | Variation (live) |
| Natural (EVA/dbSNP) variants overlapping a locus's genomic span via Ensembl |
29 | Variation (live) |
| Ensembl VEP consequence prediction for a variant (region + allele, not locus) — most-severe consequence + per-transcript SO terms, IMPACT, SIFT. All 12 organisms. |
30 | Orthology (live) |
| PANTHER protein family + subfamily (id + name), GO terms by aspect, protein class, and pathways. found=false when unmapped. All 12 organisms. |
31 | Orthology (live) |
| OrthoDB ortholog group (name, evolutionary rate) + cross-species member genes at the Viridiplantae level. organism_count + truncated. All 12 organisms. |
32 | Diversity (live) |
| AraGWAS genome-wide association hits per locus — score, MAF, SNP effect, phenotype/study. Arabidopsis-only. |
33 | Diversity (live) |
| 1001 Genomes natural-variation SNP effects across 1135 accessions — chr, position, effect, impact, amino-acid change, transcript + gene span. Arabidopsis-only. |
34 | Batch (live) |
| Parallel per-locus fanout for tools 1–6, 8–12, 14. Up to 50 loci per call. |
35 | Batch (live) |
| Runs any tool whose only required argument is |
36 | Synthesis (live) |
| Compose 2–5 backends in parallel, return a |
37 | Synthesis (live) |
| One-shot "tell me about this gene" dossier — annotation + xrefs + protein + domains + GO + KEGG + STRING + literature composed into a rendered Markdown |
38 | Families (live) |
| Every protein in one organism carrying an InterPro / Pfam / PANTHER entry, with the locus each maps to (entry → genes; the reverse of tools 23 and 30). UniProt |
39 | Families (live) |
| Every gene in an Ensembl Compara gene tree — the |
40 | Homology (live) |
| Paralogues Ensembl Compara records for a locus — |
41 | Region query (live) |
| An organism's Ensembl assembly — name, GCA accession, karyotype — and every top-level seq-region with its length: the |
42 | Provenance (live) |
| The release a backend's own endpoint calls current (Ensembl, STRING, QuickGO, JASPAR, KEGG), for the backends whose answers state none; PDBe, AraGWAS and Europe PMC publish none and say why. A separate request, so read it before and after a run: equal values mean no release changed. |
⚡ Quickstart
After install, the simplest call returns the Ensembl Plants record for
NAC001 — the canonical worked example used throughout examples/:
// arguments
{ "locus": "AT1G01010" }
// result (truncated)
{
"id": "AT1G01010",
"organism": "arabidopsis_thaliana",
"display_name": "NAC001",
"biotype": "protein_coding",
"seq_region_name": "1",
"start": 3631,
"end": 5899,
"strand": 1,
"assembly_name": "TAIR10",
"description": "NAC domain containing protein 1 ..."
}Cross-species — pass organism=:
{ "locus": "Os01g0100100", "organism": "oryza_sativa" }A recorded Claude Code session (2026-05-24) with a narrower prompt — the Ensembl record, UniProtKB entry and top three Europe PMC papers for AT1G01010 — answered it in one turn:
Full per-tool walkthroughs (with real upstream-API transcripts) live in
examples/:
Walkthrough | Coverage |
One-shot Markdown gene dossier — 7 backends composed, with graceful KEGG degradation. | |
Ensembl → xrefs → UniProt → Europe PMC → QuickGO chain (5 tools). | |
BLAST + per-hit UniProt enrichment. | |
Gramene + KEGG + UniProt + STRING + ATTED-II (5 tools). | |
The four v0.8 synthesis tools ( | |
v0.9 multi-organism resolver against rice + maize — per-backend routing. Captured 2026-05-24 against PyPI v1.0.4. |
📚 Resources & prompts
Clients discover them via resources/list and prompts/list.
Resources (resources/read):
URI | What |
| Per-backend |
| Slug → Phytozome |
| Static per-backend roster — |
| Markdown table of all 12 supported plants × 9 ID slots (ncbi_taxid / ensembl / phytozome / string / europe_pmc / kegg / atted / gprofiler / plantcyc). |
Prompts (prompts/get):
Name | Required | Optional | Chains |
|
|
| Ensembl → xrefs → UniProt → Europe PMC → QuickGO. |
|
|
|
|
|
|
| Gramene → KEGG → UniProt → STRING → ATTED-II. |
🔌 Transports
Transport | How to launch |
stdio (default) |
|
streamable-HTTP |
|
The HTTP transport is stateless and emits JSON responses by default — the right shape for registry indexers and remote hosting.
Self-hosting
There is no public hosted endpoint. To use the HTTP transport, run it
yourself: it is the same binary, gated by your own bearer token
(PLANT_GENOMICS_MCP_HTTP_TOKEN), with NCBI BLAST requests sent under your
own contact email.
⚙️ Configuration
Stdio needs no configuration. The two env vars that matter:
Variable | When | Effect |
| HTTP transport only | Bearer token for |
| If you use BLAST | NCBI etiquette contact. Unset → placeholder + per-call warning; NCBI may throttle. |
Variable | Default | Effect |
|
| HTTP bind address. |
|
| HTTP TCP port. |
|
| Reject POSTs with |
|
|
|
|
|
|
|
| Max in-flight BLAST searches per process (NCBI per-IP rate limit). |
|
| Per-backend TTL+LRU cache entry lifetime, in seconds. 200-only. |
|
| Max entries per backend before LRU eviction. |
| unset | Any non-empty value makes every cache a no-op. |
The cache is process-local — restart the server to drop all entries.
Long-running calls (retry storms, multi-second Phytozome BioMart POSTs)
emit MCP notifications/progress over the active session; clients opt
in via progressToken in the request _meta.
⚠️ Error model
All live tools raise PlantGenomicsError subclasses; the MCP SDK
stringifies them into the wire content with a [ClassName] prefix so
clients can route on failure kind without parsing the message:
Wire prefix | When |
| 404 / empty BioMart row / invalid locus identifier |
| 429 retry budget exhausted — back off and retry |
| 5xx past retry budget — service outage, try a peer backend |
| Other (BioMart |
Batch tools return {tool, count, results, errors} where
results[locus] is the same shape as the single-locus tool and
errors[locus] is the same [ClassName] message string. Ensembl's
batch uses the native POST /lookup/id endpoint (one HTTP round-trip);
everything else fans out via asyncio.gather.
🧪 Development
.venv/bin/pip install -e '.[dev]' # or: uv sync --extra dev
.venv/bin/pytest -q # unit tests
PLANT_GENOMICS_MCP_LIVE=1 .venv/bin/pytest -q # adds live network probes
PLANT_GENOMICS_MCP_STDIO_SMOKE=1 .venv/bin/pytest -q # adds stdio smoke
.venv/bin/ruff check .With uv, pass --extra dev — a bare uv sync omits (and removes) the test
dependencies. See CONTRIBUTING.md.
CI runs the unit suite + the stdio smoke on every push/PR (matrix:
Python 3.11, 3.12, 3.13, 3.14 — the full requires-python range), with a
dead proxy set so an ungated network call fails. A separate live-smoke job
runs two live test files (tests/test_verify_genes.py,
tests/test_arf_mcp_client.py, InterPro and PANTHER) on Python 3.12 on the
same pushes and PRs; it reports a result but is not a required check, because its result depends on
third-party services. The rest of the live suite (PLANT_GENOMICS_MCP_LIVE=1)
is not run in CI.
Drift detection. scripts/benchmark_annotations.py runs a curated corpus
of 27 loci (spanning all 12 organisms) through 12 functions — the organism
resolver plus 11 lookups across 9 of the 23 backends (ATTED-II, BAR, Ensembl
Plants, Europe PMC, Gramene, KEGG, Phytozome, STRING-DB, UniProt) — and
compares the results to a frozen snapshot of earlier results
(scripts/benchmark_annotations.expected.json), emitting PASS / DRIFT / FAIL
plus two cross-source consistency invariants. A change is detected relative to
that snapshot, not checked against an independent truth. BLAST and the
synthesis pipelines are registered in the script, but no corpus locus
exercises them. A scheduled GitHub Actions workflow
(.github/workflows/benchmark.yml) runs it weekly and pages when the same loci
fail on a re-run. Operator guide: docs/benchmarking.md.
.venv/bin/python scripts/benchmark_annotations.py # full live sweepSee CHANGELOG.md for release notes, including the
v0.8 → v0.9 species=/organism_id= → organism= migration and the
v1.0.1 HTTP-token enforcement change.
MCP registry
Listed in the official MCP registry
under the namespace below (ownership-verification token for mcp-publisher):
mcp-name: io.github.musharna/plant-genomics-mcpLicense
MIT — see LICENSE. Underlying services (Ensembl Plants,
Phytozome, TAIR, PlantCyc, BAR) have their own terms of use; consult
each before bulk querying.
Available Tools
56 toolsalphafold_structureAlphaFold: Predicted StructureARead-onlyIdempotent
Fetch the AlphaFold DB predicted-structure summary for a locus (alphafold.ebi.ac.uk; free, no key). Resolves the locus → UniProt accession, then returns the predicted model's global mean pLDDT confidence, the per-band pLDDT distribution with each band's pLDDT range (plddt_band_ranges), modelled residue span, latest model version, and mmCIF / PDB / PAE download URLs — links for the client to fetch; no tool on this server retrieves them. A valid protein with no deposited model returns found=false (a normal outcome, not an error); a locus with no UniProt entry raises a typed NotFoundError. Works for all 12 organisms (UniProt-keyed). Complements resolve_locus_to_uniprot (sequence-level) with the structure-level view. Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT4G09760 (Arabidopsis), Os01g0100100 (rice RAP-DB) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| gene | No | Gene name from UniProt |
| found | Yes | True if a predicted model exists |
| locus | Yes | |
| cif_url | No | mmCIF model download URL — a link for the client to open or fetch; no tool on this server dereferences it |
| pdb_url | No | PDB model download URL — a link for the client to open or fetch; no tool on this server dereferences it |
| organism | No | Organism scientific name |
| accession | Yes | Resolved UniProt accession |
| mean_plddt | No | Global mean pLDDT confidence (0–100) |
| description | No | UniProt protein description |
| plddt_bands | No | Fraction of residues per confidence band |
| model_created | No | Model creation date (ISO 8601) |
| pae_image_url | No | Predicted-aligned-error image URL — a link for the client to open or fetch; no tool on this server dereferences it |
| residue_range | No | Modelled residue span {start, end} |
| latest_version | No | Latest AlphaFold model version |
| model_entity_id | No | e.g. AF-Q9SZ92-F1 |
| upstream_version | No | AlphaFold DB release that produced THIS response, as stated by the entry's own latestVersion (e.g. '6'). null means AlphaFold DB did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| plddt_band_ranges | Yes | Band -> [lower, upper] pLDDT on the 0-100 scale, per EMBL-EBI: very_low <50, low 50-70, confident 70-90, very_high >90 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only, idempotent, and non-destructive, but the description adds important behavioral details: it resolves locus to UniProt, returns specific fields (plddt confidence, band ranges, model version, download URLs), and defines edge cases—found=false for valid proteins without a model (a normal outcome) and a typed NotFoundError for missing UniProt entries. This goes well beyond the annotations and gives an agent a clear picture of outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized: it opens with the core purpose, enumerates the returned data, covers edge cases and errors, notes the sibling relationship, and closes with defaults. Every sentence contributes value; no filler. It front-loads the primary action and resource, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters, an output schema, and several nuanced behaviors (resolution, error types, default organism, links-to-fetch-only), the description covers all of them. It even mentions the full organism set (12 organisms, UniProt-keyed). Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters already have descriptive text (e.g., locus examples, organism acceptors). The description reinforces the default organism and explains the locus resolution, but it does not add significant meaning beyond the schema. Baseline 3 is appropriate given the schema's completeness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Fetch the AlphaFold DB predicted-structure summary for a locus', naming the external service and the input type. It also differentiates from the sibling resolve_locus_to_uniprot by explicitly stating this is the structure-level view, so an agent can immediately tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the sibling tool and the distinction ('Complements resolve_locus_to_uniprot (sequence-level) with the structure-level view'). It also gives practical usage context: defaults to arabidopsis_thaliana, the need to pass organism for others, and that the returned URLs are for the client to fetch (not for the tool to download). These are direct when/how-to-use cues.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_locus_synthSynthesis: Locus OverviewARead-onlyIdempotent
Synthesis: one-call equivalent of the analyze_locus prompt. Resolves a locus through Ensembl Plants, then fans out to xrefs, UniProt, Europe PMC, and QuickGO in parallel. Returns a SynthesisEnvelope with per-step status and a reconciled summary flagging cross-source name/accession disagreements.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Locus name, e.g. AT1G01010 | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | Synthesis tool name, e.g. analyze_locus_synth |
| input | Yes | Echoed input arguments |
| steps | Yes | Per-backend execution rows |
| result | No | Composed cross-source result; None if root step failed |
| elapsed_s | Yes | Total orchestrator wall time |
| started_at | Yes | ISO 8601 UTC timestamp |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety bar is met. The description adds genuinely useful behavior beyond annotations: parallel fan-out, per-step status reporting, and the reconciled summary that flags cross-source name/accession disagreements. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler, front-loaded with the synthesis purpose and immediately followed by scope and return details. The phrase 'one-call equivalent of the analyze_locus prompt' is slightly redundant with the synthesis framing but still earns its place as context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, yet the description still names the SynthesisEnvelope and its contents. Given the orchestration complexity, the description covers what it does, which sources it touches, and what it returns; the missing piece is explicit routing guidance versus the individual sibling tools, but annotations and schema cover the rest.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with full descriptions for both locus and organism parameters, so the baseline is 3. The description mentions resolving 'a locus' but adds no parameter-specific syntax, format, or behavior details beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (resolves) and resource (locus via Ensembl Plants), then names the exact fan-out targets (xrefs, UniProt, Europe PMC, QuickGO). The 'Synthesis: one-call equivalent' framing clearly differentiates it from the many individual sibling tools like get_gene_xrefs or resolve_locus_to_uniprot, telling the agent this is the aggregated variant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'One-call equivalent of the analyze_locus prompt' implies it is the all-in-one alternative to calling individual sibling tools separately, but the description never explicitly says when to prefer it versus calling get_gene_xrefs or resolve_locus_to_uniprot alone, nor states exclusions. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arabidopsis_natural_variation1001 Genomes: Natural VariationARead-onlyIdempotent
Fetch 1001 Genomes natural-variation SNP effects for an Arabidopsis locus (tools.1001genomes.org; free, no key) — the variation observed across 1135 resequenced natural accessions. Returns per-SNP effect rows (chromosome, position, accession id, effect, impact, amino-acid change, transcript) plus the gene's genomic span. variant_count is the true row total even when capped. ARABIDOPSIS-ONLY — any other organism raises OrganismNotSupported. Defaults to arabidopsis_thaliana.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Arabidopsis AGI locus, e.g. AT1G01060 (a bare AGI is transcript-scoped to .1) | |
| cursor | No | next_cursor from the previous page; omit for the first (#123) | |
| organism | No | Arabidopsis only (the 1001 Genomes panel is A. thaliana) | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True once the effects endpoint returned 200 |
| locus | Yes | |
| total | Yes | How many variant effects exist upstream for this query, all pages (pre-cap) (#123) |
| region | No | Genomic span, e.g. 'Chr1:33666..37840' |
| organism | Yes | Always arabidopsis_thaliana |
| returned | Yes | Effect rows returned (post-cap) |
| variants | No | Per-effect {chr, position, accession_id, effect, impact, amino_acid_change, …} |
| truncated | Yes | True if the effect list was capped |
| transcript | Yes | Transcript-scoped gene id used (e.g. AT1G01060.1) |
| next_cursor | No | Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123) |
| variant_count | Yes | Total effect rows (pre-cap) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, and non-destructive behavior, and the description adds valuable context beyond that: the service is free and requires no key, pagination has a true-count caveat ('variant_count is the true row total even when capped'), and unsupported organisms raise a specific error. This helps an agent reason about external-service behavior, truncation, and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences with the core action front-loaded, followed by return content and critical caveats. Parentheticals add relevant detail without excessive padding. It could be slightly tighter, but every sentence carries useful information and the structure is logical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, so the description correctly focuses on selection criteria, organism restriction, pagination semantics, and external-service access details. An agent has enough information to decide whether to call this tool and to set expectations about errors and result limits. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters, including the AT1G01060 example, cursor usage, and organism default. The description repeats the default organism but does not add substantive parameter-level meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Fetch 1001 Genomes natural-variation SNP effects for an Arabidopsis locus.' It clearly defines the data scope (1135 resequenced accessions) and the tool's specialized output, setting it apart from generic locus or variant tools. It also states the return payload concretely, so an agent knows exactly what this tool produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong usage context by naming the data source, explaining the organism restriction ('ARABIDOPSIS-ONLY'), and stating the default organism. It also warns that non-Arabidopsis organisms raise OrganismNotSupported. However, it does not explicitly compare this tool to sibling alternatives like locus_variants, so the when-to-use-vs-alternatives guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aragwas_associationsAraGWAS: GWAS AssociationsARead-onlyIdempotent
Fetch AraGWAS genome-wide association study hits for an Arabidopsis locus (aragwas.1001genomes.org; free, no key). Returns each significant SNP association overlapping the gene with its score (-log10 p), minor-allele frequency, the SNP's predicted molecular effect (impact, amino-acid change), and the phenotype/study it came from, including the study's own significance thresholds on the score's scale (study.thresholds), which the over_bonferroni / over_fdr / over_permutation flags are taken against. Rows come strongest first, 25 per answer by default (limit, up to 100); next_cursor resumes at the first row not returned. association_count is the true total even when page-capped. ARABIDOPSIS-ONLY — any other organism raises OrganismNotSupported. Defaults to arabidopsis_thaliana.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max associations, strongest first (1–100, default 25) | |
| locus | Yes | Arabidopsis AGI locus, e.g. AT1G01060 | |
| cursor | No | next_cursor from the previous page; omit for the first (#123) | |
| organism | No | Arabidopsis only (the 1001 Genomes panel is A. thaliana) | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True once the associations endpoint returned 200 |
| locus | Yes | |
| total | Yes | How many associations exist upstream for this query, all pages (pre-cap) (#123) |
| organism | Yes | Always arabidopsis_thaliana |
| returned | Yes | Associations returned (post page-cap) |
| truncated | Yes | True if pagination was capped |
| next_cursor | No | Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123) |
| associations | No | Per-hit {score, maf, mac, over_bonferroni, over_fdr, over_permutation, snp{…}, study{…, thresholds}}. score is -log10(p); study.thresholds holds that study's bonferroni_threshold05/01, bh_threshold and permutation_threshold on the same scale, which over_bonferroni / over_fdr / over_permutation compare score against |
| upstream_version | No | AraGWAS release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='aragwas') reports the release its own endpoint calls current at query time, or why there is none. null means AraGWAS did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| association_count | Yes | Total associations (pre-cap) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly/idempotent/destructive=false), and the description goes well beyond them: it discloses no-auth access, deterministic ordering (strongest first), page size and cap, cursor semantics, the guarantee that association_count is the true total even when page-capped, and the OrganismNotSupported failure mode. That is substantial operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph with the core action front-loaded and pagination/organism constraints at the end. Nearly every clause carries unique information, though it is long enough that mild trimming would improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, output-schema-backed tool, the description covers everything an agent needs: what is returned, ordering, pagination, the true-total caveat, authentication needs, and the error condition. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: the ordering guarantee behind limit (strongest first), the resume semantics of next_cursor (first row not returned), and the organism constraint ('Arabidopsis only'). This raises it above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (AraGWAS GWAS association hits) scoped to an Arabidopsis locus, and enumerates the returned fields (score, MAF, predicted effect, phenotype/study, thresholds). This clearly distinguishes it from neighbors like arabidopsis_natural_variation and locus_variants without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its use case (pull GWAS hits for a locus) and notes it is free/no key, plus a hard constraint that non-Arabidopsis input raises OrganismNotSupported. However it never routes the agent among the many sibling locus tools (e.g., arabidopsis_natural_variation, locus_variants) or states when this source is preferable, leaving usage to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
atted_coexpressionATTED-II: CoexpressionARead-onlyIdempotent
Fetch co-expressed gene neighbors from ATTED-II (atted.jp, API v5) for a plant locus. Returns top_n neighbors with target locus + NCBI Entrez gene ID + score (higher = stronger coexpression), in the index the release declares as score_type: 'z' for Ath-u.c4-0, 'LSmr' (logit score) for the other releases. The ATTED-II release (e.g. Ath-u.c4-0 for Arabidopsis, Osa-u.c1-0 for rice) is resolved per-organism. Covers: arabidopsis_thaliana, glycine_max, medicago_truncatula, oryza_sativa, solanum_lycopersicum, vitis_vinifera, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms). A locus that is not in the organism's release raises NotFoundError. Pairs with string_interactions to surface high-confidence functional partners (interactors that are also coexpressed).
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Plant locus, e.g. AT1G01010 (Arabidopsis) or Os01g0100100 (rice) | |
| top_n | No | ||
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| locus | Yes | |
| total | No | Always null: the upstream returns the top coexpression neighbours asked for and states no total; null means unknown, never zero (#123) |
| returned | Yes | Rows in this payload (#123) |
| neighbors | Yes | |
| truncated | No | Always null: without a stated total, truncation is unknown (#123) |
| score_type | Yes | The coexpression index every neighbour's score is in, as ATTED-II declares it: 'z' for Ath-u.c4-0, 'LSmr' (logit score) for the other releases |
| atted_release | Yes | ATTED-II DB identifier, e.g. Ath-u.c4-0 (release version included) |
| upstream_version | No | ATTED-II release that produced THIS response, as stated by the db= pinned in the request (e.g. 'Ath-u.c4-0'); same value as atted_release. null means ATTED-II did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, it discloses concrete behavior: per-organism release resolution, explicit error behavior (OrganismNotSupported before any request, NotFoundError for missing loci), score_type semantics, and the organism allowlist. This is substantial context an agent would not get from the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and output, followed by release, organism, and error details in a logical order. The final 'Pairs with string_interactions' sentence is ambiguous and slightly undercuts the otherwise tight structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and annotations already exist, the description is largely complete: it covers organisms, errors, scoring, and release handling. Minor unresolved details are the exact meaning of 'Pairs with string_interactions' and possible API limits, which are not critical for calling the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and locus/organism already have decent descriptions. The description adds value by clarifying top_n as the number of returned neighbors, enumerating the organisms accepted by organism, and explaining release resolution. It does not merely repeat the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch co-expressed gene neighbors from ATTED-II.' It names API v5, the input (plant locus), the output shape (top_n neighbors with locus, Entrez ID, score), and the supported organisms, so an agent can distinguish it from sibling interaction tools like string_interactions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states what it returns and lists supported organisms, but it never explicitly says when to choose this over coexpression or interaction alternatives, and it does not give exclusions. The final sentence about pairing with string_interactions is too vague to count as solid usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bar_aiv_interactionsBAR: Predicted InteractionsARead-onlyIdempotent
Fetch BAR AIV (Arabidopsis Interactions Viewer) interactions for an Arabidopsis or rice locus. Dispatches by organism: Arabidopsis returns curated GRN paper refs from /interactions/get_paper_by_agi/{locus} (PubMed ID, title, image URL, comments, pipe-split tags; the image is a link for the client to fetch, no tool on this server retrieves it); rice returns predicted PPI partners from /interactions/rice/{locus} with Pearson co-expression r (pcc), evidence hits, and quality score. The kind field discriminates the response shape (grn_papers vs ppi_predictions). Rice requires the MSU LOC_Os* locus format — RAP-DB Osg is rejected upstream. Only Arabidopsis and rice are supported by AIV; other organisms raise OrganismNotSupported.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | AGI locus (AT1G01010) for Arabidopsis or MSU locus (LOC_Os01g01080) for rice | |
| organism | No | arabidopsis_thaliana or oryza_sativa — slug, scientific/common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| kind | Yes | Discriminator: grn_papers (Arabidopsis) or ppi_predictions (rice) |
| count | Yes | Total rows returned (len of papers or partners) |
| locus | Yes | |
| papers | No | GRN paper refs (populated when kind=grn_papers) |
| organism | Yes | |
| partners | No | PPI predictions (populated when kind=ppi_predictions) |
| source_url | Yes | BAR AIV endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and idempotentHint annotations, the description discloses substantive behavior: the response shape is discriminated by the `kind` field, the image URL is only a link and no tool on this server retrieves it, rice rejects RAP-DB Os*g* format upstream, and unsupported organisms raise OrganismNotSupported. This is rich, accurate behavioral context that goes well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It front-loads the core purpose, then efficiently covers endpoint routing, response fields, organism-specific caveats, and failure behavior. There is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given this tool has an output schema and complete parameter schema coverage, the description adds the remaining context an agent needs: supported organisms, endpoint dispatch, response shape discrimination, the MSU locus requirement, and error behavior. Nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema already documents both parameters well (100% coverage), the description adds essential semantics beyond it: the exact locus formats expected for each organism, the rejection of RAP-DB identifiers for rice, and how the `kind` field in the response relates to the organism-specific behavior. These details materially help an agent select correct parameter values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch BAR AIV (Arabidopsis Interactions Viewer) interactions for an Arabidopsis or rice locus.' It clearly defines the tool's scope, distinguishes it from generic interaction tools by naming AIV and the unique endpoints, and calls out the supported organisms. The purpose is unambiguous and not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use the tool: for Arabidopsis or rice AIV interactions, and explicitly says other organisms raise OrganismNotSupported. It also gives decisive routing guidance (curated GRN papers for Arabidopsis, predicted PPI partners for rice) and locus-format requirements. However, it does not explicitly name alternative tools or state when a sibling should be preferred, so it falls just short of full usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bar_efp_expressionBAR: eFP ExpressionARead-onlyIdempotent
Fetch BAR/eFP world-map natural-variation expression for an Arabidopsis locus. Wraps the world-eFP view at /microarray_gene_expression/world_efp/arabidopsis/{locus} — returns expression across ~36 ecotypes (Bay-0, Col-0, Cvi-1, Ler-2, ...) with per-replicate values, control samples, collection lat/lng, and a per-ecotype mean computed client-side. Arabidopsis only. BAR is keyless and a Global Core Biodata Resource (2023).
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Arabidopsis AGI locus, e.g. AT1G01010 |
Output Schema
| Name | Required | Description |
|---|---|---|
| locus | Yes | |
| species | Yes | |
| ecotypes | Yes | |
| probeset | No | Microarray probeset ID, uniform across ecotypes for one gene |
| source_url | Yes | BAR world-eFP endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it |
| ecotype_count | Yes | Number of ecotype rows in `ecotypes` |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and non-destructive. The description adds valuable behavioral details: the endpoint it wraps, the fields returned (per-replicate, control samples, lat/lng), the client-side computation of the mean, and the approximate ecotype count. It also notes keyless access and the resource's status. No contradictions; it complements the annotations well.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized but packs essential information: main purpose, endpoint, output fields, ecotype count, keyless access, and resource status. The most critical information (what it fetches) is front-loaded, and every sentence adds value. It could be slightly tighter, but it avoids redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema (not shown here but indicated), the description needn't detail the full return structure, but it does list the main fields. It covers constraints (Arabidopsis only, keyless), the nature of the data, and the source. For a simple one-parameter read-only tool, this is complete enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single parameter (locus), with a clear AGI format example. The description adds minimal additional meaning beyond the schema—it reiterates the locus usage and Arabidopsis restriction but doesn't introduce new syntax or formatting guidance. Baseline 3 is appropriate because the schema already fully documents the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Fetch'), a precise resource (BAR/eFP world-map natural-variation expression), and the object (Arabidopsis locus). It distinguishes itself from siblings by specifying 'world-map natural-variation expression' and the exact endpoint path, leaving no ambiguity about what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context: this is for Arabidopsis natural-variation expression across ecotypes, and notes 'Arabidopsis only' and 'BAR is keyless.' While it doesn't explicitly name alternative tools or when not to use it, the specificity of the domain (expression vs. summary, interactions, etc.) implicitly guides selection. A minor gap is the absence of explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bar_gene_summaryBAR: Gene SummaryARead-onlyIdempotent
Fetch the BAR (Bio-Analytic Resource, U Toronto) merged ThaleMine + GAIA-aliases summary for an Arabidopsis locus. Returns the TAIR curator summary + Araport11 computational description from /thalemine/gene_information/ together with the NCBI Gene ID and cross-DB aliases (RefSeq, UniProt, TIGR locus-model IDs) from /gaia/aliases/. Arabidopsis only — ThaleMine carries taxon 3702 plus yeast/human for ortholog cross-reference. BAR is keyless and a Global Core Biodata Resource (2023); replaces the v0.9 subscription-gated tair_locus_info stub for the curator-summary use case.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Arabidopsis AGI locus, e.g. AT1G01010 |
Output Schema
| Name | Required | Description |
|---|---|---|
| agi | No | AGI primary identifier echoed by ThaleMine, e.g. "AT1G01010" |
| locus | Yes | |
| symbol | No | Gene symbol, e.g. "NAC001" |
| aliases | No | Cross-DB aliases from /gaia/aliases/ (RefSeq accessions, UniProt accessions, TIGR locus-model IDs, and TAIR aliases). Empty list if /gaia degraded. |
| species | Yes | |
| synonyms | No | TAIR aliases (CSV from Gene.tairAliases, split on commas + stripped) |
| full_name | No | Gene name from ThaleMine |
| source_url | Yes | ThaleMine endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it |
| ncbi_gene_id | No | NCBI Gene ID from /gaia/aliases/ — None if BAR has no NCBI cross-ref |
| tair_locus_id | No | TAIR locus ID from Gene.secondaryIdentifier, e.g. "locus:2200935" |
| curator_summary | No | Gene.tairCuratorSummary — the TAIR-curated functional summary prose |
| brief_description | No | Gene.briefDescription — short blurb (often same as full_name) |
| tair_short_description | No | Gene.tairShortDescription — TAIR-specific short description |
| computational_description | No | Gene.tairComputationalDescription — Araport11-sourced computed description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description adds valuable behavioral context: BAR is keyless, no subscription is needed, data comes from specific endpoints, and the underlying resource carries taxon 3702 plus yeast/human for ortholog cross-reference. This meaningfully supplements the readOnly/idempotent annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and return contents, and is dense but not bloated. The phrase 'Global Core Biodata Resource (2023)' is a credibility credential rather than operationally necessary detail, and 'keyless' plus the endpoint breakdown are the most valuable extras.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema, the description covers the input scope, data provenance, exact endpoints, returned fields, access model, and the intended replacement use case. Nothing essential is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The lone parameter 'locus' is already fully documented in the schema with the same Arabidopsis AGI locus format and example (AT1G01010). The description does not add new precision about accepted formats or edge cases, so the schema carries the weight here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Fetch the BAR ... merged ThaleMine + GAIA-aliases summary for an Arabidopsis locus.' It enumerates the returned content (TAIR curator summary, Araport11 description, NCBI Gene ID, cross-DB aliases) and explicitly distinguishes itself from the older tair_locus_info stub, making its role clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear use context: it is for Arabidopsis locus curator summaries, excludes non-Arabidopsis use, and explicitly positions itself as replacing the subscription-gated tair_locus_info stub. It does not, however, contrast itself against other sibling locus/xref tools such as get_gene_xrefs or ensembl_plants_lookup_locus.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_atted_coexpressionBatch: ATTED-II CoexpressionARead-onlyIdempotent
Batch version of atted_coexpression. Up to 50 loci per call. Covers: arabidopsis_thaliana, glycine_max, medicago_truncatula, oryza_sativa, solanum_lycopersicum, vitis_vinifera, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | ||
| top_n | No | ||
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, open-world, idempotent, and non-destructive behavior, so the bar is lower. The description adds concrete behavioral details: the 50-locus cap and early OrganismNotSupported error for unsupported organisms, which go beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler. Key constraints are front-loaded, and each sentence adds useful information: batch nature, capacity, supported organisms, and error behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and simple wrapper semantics, the description covers batch capacity, supported organisms, and failure mode. It could more explicitly state when to choose this over the single atted_coexpression, but the batch naming and sibling context make that reasonably clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description must compensate. It adds the supported organism list and the 50-locus cap, but the cap is already in schema maxItems and it says nothing about top_n, leaving a core parameter's meaning undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States it is the batch version of atted_coexpression, with a concrete batch limit of 50 loci and a list of supported organisms. This distinguishes it from the single-locus sibling, though it never explicitly states the action (e.g., 'retrieves coexpression data').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit batch capacity (up to 50 loci), enumerates supported organisms, and states that unsupported organisms raise OrganismNotSupported before any request. It implies use for batch queries but does not explicitly contrast with the single atted_coexpression tool or other batch siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_bar_aiv_interactionsBatch: BAR Predicted InteractionsARead-onlyIdempotent
Batch variant of bar_aiv_interactions. Fans out per-locus BAR AIV calls in parallel (up to 50 loci); all loci in a single call share the same organism. Each results[locus] is the full single-locus payload (kind=grn_papers for Arabidopsis with papers list, kind=ppi_predictions for rice with partners list).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. | |
| organism | No | arabidopsis_thaliana or oryza_sativa — slug, scientific/common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds context beyond annotations: parallel execution, 50-locus limit, organism constraint per call, and the varying result payload per organism (papers vs partners). It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences efficiently convey purpose, constraints, and result structure without redundancy. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the batch complexity and existing output schema, the description covers parallelism, limit, organism sharing, and result structure. Minor gap: no explicit error handling beyond mentioning errors[locus], but overall sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds that loci are identifiers with success/failure mapping in results and errors, but does not significantly extend schema-provided parameter meanings.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a batch variant of bar_aiv_interactions, outlines the parallel fan-out for up to 50 loci, and specifies the result structure per locus. It distinguishes itself from the single-locus sibling by focusing on batch processing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when querying multiple loci, noting the constraint that all loci share the same organism. It distinguishes from the single-locus variant but does not explicitly compare with other batch tools or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_bar_gene_summaryBatch: BAR Gene SummaryARead-onlyIdempotent
Batch variant of bar_gene_summary. Fans out per-locus BAR ThaleMine + GAIA-aliases calls in parallel (up to 50 loci). Each results[locus] is the full single-locus payload (curator summary, computational description, NCBI Gene ID, cross-DB aliases). Arabidopsis only.
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false. Description adds behavioral context: fan-out in parallel, up to 50 loci, result structure, and organism restriction. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and constraints. No extraneous information. Every sentence adds value (batch variant, parallel fan-out, payload summary, organism restriction).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists (as per context), the description sufficiently explains return values without needing full schema details. Covers key payload fields and error handling, and includes scope (Arabidopsis) and limits (50 loci).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a description for the loci parameter. Description adds meaning about array items and error handling ('Successes land in results[locus]; PlantGenomicsError failures in errors[locus]'), which goes beyond the schema-provided description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it's a batch variant of bar_gene_summary, describes parallel fan-out over up to 50 loci, and specifies the output structure (curator summary, computational description, NCBI Gene ID, cross-DB aliases). Distinguishes from sibling bar_gene_summary by being batch and from other batch tools by specifying the exact per-locus payload.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Batch variant of bar_gene_summary' and 'Arabidopsis only', implying use for multiple loci. Does not explicitly state when not to use (e.g., for single locus), but the context is clear. Provides information about parallel execution and payload structure.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_ensembl_plants_lookup_locusBatch: Ensembl Plants Locus MetadataARead-onlyIdempotent
Batch variant of ensembl_plants_lookup_locus. Uses Ensembl's native POST /lookup/id endpoint — one HTTP round-trip for up to 50 loci, materially cheaper than N parallel GETs. Successes in results[] with the same shape as the single-locus tool. Retries 429/5xx via the shared _http helper (Retry-After capped at 60 s). Misses (loci with no record) still land in errors[] with the [NotFoundError] prefix; the whole batch only fails when the retry budget is exhausted.
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses retry behavior (429/5xx via _http helper, Retry-After capped at 60s), explains error handling (misses in errors[] with prefix), and notes result shape matches single-locus tool. Adds value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, no redundancy. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given batch complexity, error handling, and cost benefit, the description is thorough. Output schema exists, annotations cover safety, so all behavioral aspects are addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline 3. Description adds that loci max is 50 and clarifies error structure, providing context beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it's a batch variant of a single-locus tool, uses POST /lookup/id endpoint for up to 50 loci, and highlights cost efficiency. Distinguishes from sibling 'ensembl_plants_lookup_locus'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly guides use for multiple loci via 'batch variant' and 'cheaper than N parallel GETs'. No explicit when-not or alternative mentions, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_get_gene_xrefsBatch: Gene Cross-ReferencesARead-onlyIdempotent
Batch variant of get_gene_xrefs. Fans out per-locus xref lookups over Ensembl Plants in parallel (up to 50 loci). Each results[locus] is the full single-locus shape (count + xrefs[] + by_db rollup).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, open-world, non-destructive behavior. The description adds valuable context: parallel fan-out, 50-locus limit, and the output structure per locus. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. It front-loads the core identity ('Batch variant of get_gene_xrefs') and then details behavior concisely.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of full annotations and an output schema, the description is complete. It covers the batch nature, parallelism, limits, and output shape. No critical missing information for effective tool use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description does not add meaning beyond the schema; it mentions the loci parameter implicitly but does not elaborate on organism. The parameter semantics are adequately covered by the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies this as a batch variant of get_gene_xrefs, describes the parallel fan-out over Ensembl Plants with a max of 50 loci, and specifies the output shape per locus. It effectively distinguishes this tool from the single-locus version.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly suggests use for multiple loci, but lacks explicit guidance on when to use this vs alternatives (e.g., the single-locus get_gene_xrefs or other batch tools). No when-not-to-use or alternative mentions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_gramene_homologsBatch: Gramene HomologsARead-onlyIdempotent
Batch version of gramene_homologs. Up to 50 loci per call; shares the homology_type filter across all loci. Returns the standard batch envelope (count + results dict + errors dict).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (max 50) | |
| homology_type | No | ortholog | |
| with_organism | No | Add 'organism' (Gramene species slug, null when unknown) to every row without filtering; one extra call per 100 rows (#130) | |
| target_organism | No | Keep only homologs in this organism (slug, scientific/common name, or NCBI taxid), filtered BEFORE the cap so a hub gene's rice or wheat orthologs cannot be pushed past 'limit' by other species. Adds 'organism' to every row and 'total_all_organisms'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and idempotent safety. The description adds the 50-locus limit, the shared filter behavior, and the return envelope structure (count + results + errors), which are useful behavioral details beyond the schema and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. It front-loads the batch nature and key constraints (cap, filter sharing, return format) and avoids unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and detailed parameter descriptions, the description covers the essential cap and filter behavior. It doesn't mention optional parameters like with_organism or target_organism, but those are fully described in the schema, so an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75% (homology_type lacks a description). The description adds that homology_type is shared across all loci, giving meaning to that parameter. Other parameters are well-documented in the schema, so the description adds value where needed without redundancy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a batch version of gramene_homologs, with a 50-locus cap and shared homology_type filter. This distinguishes it from the single-locus sibling and conveys the core function without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Batch version' phrasing implies use for multiple loci, and the 50-locus cap hints at splitting larger requests. However, it doesn't explicitly name the single-locus alternative or state when not to use it, nor does it differentiate from other batch tools like batch_locus_call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_kegg_pathwaysBatch: KEGG PathwaysARead-onlyIdempotent
Batch version of kegg_pathways. Up to 50 loci per call. Covers: arabidopsis_thaliana, brachypodium_distachyon, glycine_max, hordeum_vulgare, oryza_sativa, populus_trichocarpa, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | ||
| organism | No | Plant organism — accepts canonical slug, scientific or common name, or NCBI taxid; see the tool description for which KEGG covers | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, and non-destructive behavior. The description adds valuable traits: the 50-loci cap and the pre-request validation that raises OrganismNotSupported for unsupported organisms. This goes beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, each carrying essential information: batch nature, limit, supported organisms, and error behavior. The content is front-loaded and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only batch tool with an output schema and safety annotations, the description covers the key operational details: batch limit, organism coverage, and error handling. It does not explain what the loci parameter should contain, but that may be inferred from the sibling kegg_pathways tool. Overall, the description is sufficient for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the organism parameter with a description, but the loci parameter has no schema description, and the tool description does not clarify what loci should be (e.g., gene IDs, locus identifiers). The description does add the list of supported organisms and the error behavior, but it only partially compensates for the undocumented loci parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the tool as the 'Batch version of kegg_pathways', which clearly indicates its purpose as a batched pathway lookup. It specifies the batch limit (50 loci) and enumerates supported organisms, differentiating it from the single-locus sibling and other batch tools. However, it never explicitly states that it returns KEGG pathways for the given loci, relying on the sibling's implied functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage for multiple loci via the 'batch version' phrasing and sets clear limits. It also explicitly warns that unsupported organisms raise an error before any request, guiding the agent to validate organisms upfront. It does not name alternatives or explicitly state when not to use it, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_locus_callBatch: Any Locus ToolARead-onlyIdempotent
Run one locus-keyed tool over up to 50 loci in one call (#131). 'tool' names any tool whose only required argument is 'locus' (interpro_domains, alphafold_structure, panther_family, orthodb_orthologs, gene_report, ...); 'args' holds that tool's other arguments, shared by every locus and checked against its schema once before any call. Returns the standard batch envelope: results keyed by locus, each exactly what the single tool returns, and per-locus errors. The dedicated batch_* tools remain.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | The tool's arguments other than 'locus' (and not 'cursor'), e.g. {"organism": "oryza_sativa"} | |
| loci | Yes | Locus identifiers (max 50) | |
| tool | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description substantially enriches the annotations by disclosing that args are shared across loci and validated once against the target tool's schema, that the response is a standard batch envelope keyed by locus, that each result matches the single-tool output, and that per-locus errors are included. This is exactly the behavioral nuance an agent needs beyond readOnly/idempotent hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the core action, limit, examples, behavior, and relationship to dedicated batch tools all conveyed in three sentences. The only minor drag is the issue-number reference (#131), which is noise for an agent selecting the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the annotations already carry the safety profile, the description covers what an agent needs: eligibility, argument handling, limits, and the response envelope. It could be slightly more explicit about when the generic batch path should be chosen over dedicated batch tools, but nothing required for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, so the description must compensate, and it does. It explains the meaning of 'tool' beyond the enum, defines 'args' as the target tool's non-locus arguments, and states that args are shared and schema-checked. It adds real semantic value over the bare schema, though the schema already covers the 'loci' array limit and basic args shape.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Run one locus-keyed tool over up to 50 loci in one call'. It further clarifies the exact class of tools eligible ('any tool whose only required argument is "locus"') with concrete examples, making it easy to distinguish from the single-locus tools and the dedicated batch_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the eligibility condition for the 'tool' argument and notes that 'dedicated batch_* tools remain', but it does not explicitly say when to prefer this generic batch tool over a dedicated batch_* sibling. The use case is implied (batch many loci through one locus-keyed tool) rather than directly contrasted with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_locus_go_annotationsBatch: GO AnnotationsARead-onlyIdempotent
Batch variant of locus_go_annotations. Two-stage fanout — each locus is resolved to UniProt and then queried in QuickGO. Per-locus NotFoundError from either stage lands in errors[] with the typed prefix preserved. Capped at 50 loci.
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. | |
| limit | No | Max annotations per locus from QuickGO (1–100, default 50) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds significant behavioral detail: the two-stage fanout, per-locus error handling (NotFoundError into errors[] with typed prefix), and the 50-locus cap. This goes well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences, each adding distinct information: purpose, process, error behavior, and cap. No unnecessary words; front-loaded with the most important information (batch variant).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a batch tool with a two-stage process, the description covers the key aspects. It mentions error handling and the cap. However, it could benefit from a brief note on when to prefer this over other batch tools or the single-locus version, though the tool name and sibling set provide some context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description adds context about error handling for the 'loci' parameter (how errors are structured). It also clarifies the default for 'organism' and the range for 'limit'. This adds value beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it is a batch variant of locus_go_annotations, specifies the two-stage fanout (locus to UniProt to QuickGO), and includes a cap of 50 loci. This distinguishes it from the single-locus sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for multiple loci by being a batch variant, and mentions the cap of 50 loci. However, it does not explicitly state when not to use it or compare with other batch tools (e.g., batch_locus_literature) to avoid confusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_locus_literatureBatch: LiteratureARead-onlyIdempotent
Batch variant of locus_literature. Fans out per-locus Europe PMC searches in parallel (up to 50 loci). Each results[locus] is the full single-locus payload (query + hitCount + returned + hits[]).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. | |
| size | No | Max results per locus (1–25, default 10) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds value by explaining parallel execution and error handling (successes in results, errors in errors), going beyond what annotations offer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. It front-loads the purpose ('Batch variant of locus_literature') and efficiently covers parallelism, limits, and result structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (batch, parallel, error handling), the description covers all key aspects: it's a batch, fan-out, max 50 loci, per-locus response payload. Output schema exists, so return values are covered. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the description does not add significant meaning beyond what the schema already provides. It mentions 'up to 50 loci' which is already in schema maxItems. Baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a batch variant of locus_literature, specifies parallel per-locus searches for up to 50 loci, and describes the results structure. This distinguishes it effectively from the single-locus sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when multiple loci are involved but does not explicitly state when to use this tool versus other batch tools like batch_locus_go_annotations or alternatives. No direct comparison or exclusion criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_phytozome_lookup_locusBatch: Phytozome Locus MetadataARead-onlyIdempotent
Batch variant of phytozome_lookup_locus. Fans out per-locus BioMart queries in parallel (up to 50 loci). Each results[locus] is the full single-locus row (organism_name, gene_name, chromosome, start/end/strand, description).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive. The description adds that it fans out in parallel with a 50-locus limit and specifies error handling (errors[locus]), providing useful context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: first sets context (batch, parallel, limit), second describes output structure. Every sentence earns its place with no superfluous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, the description adequately covers batch behavior, limits, and error handling. However, comparing to siblings, it could better highlight when to prefer this batch tool over others.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds value by explaining loci limits and stating that organism accepts slugs, scientific/common names, or NCBI taxids, aiding correct parameter usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this is a batch variant that fans out per-locus BioMart queries in parallel, and specifies the max loci and output structure, distinguishing it from the single-locus sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It identifies as a batch version but does not provide explicit guidance on when to use it versus other batch tools or the single-locus version, missing opportunities to guide selection among numerous siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_resolve_locus_to_uniprotBatch: Resolve Locus → UniProtARead-onlyIdempotent
Batch variant of resolve_locus_to_uniprot. Fans out per-locus UniProtKB searches in parallel (up to 50 loci). Each results[locus] is the full single-locus record (primaryAccession + uniProtkbId + entryType + geneNames + organism + sequenceLength + web_url + …).
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | List of locus identifiers (1–50). Successes land in results[locus]; PlantGenomicsError failures in errors[locus]. | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds valuable behavioral context: parallel fan-out, batch size limit of 50, and error handling per locus. This goes beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences succinctly convey the tool's batch nature, parallelism, limits, and output structure. Every sentence is informative and well-placed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and detailed annotations, the description fully covers the tool's purpose, behavior, limits, and result format. No critical gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter has a schema description. The tool description adds no additional semantic information about parameters beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a batch variant of resolve_locus_to_uniprot, performing parallel searches for up to 50 loci. It explicitly distinguishes from the single-locus sibling by specifying batch behavior and result structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates it is a batch variant, implying use for multiple loci, but does not explicitly state when not to use or list alternatives among sibling batch tools. However, the context is clear given the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_string_interactionsBatch: STRING InteractionsBRead-onlyIdempotent
Batch version of string_interactions. Up to 50 inputs per call. The argument was called loci_or_accessions before #129; that name is still accepted, deprecated.
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | ||
| limit | No | ||
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | The batch tool name, e.g. batch_resolve_locus_to_uniprot |
| count | Yes | Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus. |
| errors | Yes | locus → '[ClassName] message' for PlantGenomicsError failures |
| results | Yes | locus → per-locus result dict (same shape as the single-locus tool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds the batch limit and a deprecation note about the parameter name change, which are useful. However, it does not disclose how partial failures are handled, rate limits, or other batch-specific behaviors, leaving some gaps beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. It front-loads the key batch limit and then adds the deprecation note. It is efficient and easy to parse, though it could have included a bit more context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a batch tool with three parameters and an output schema present, the description is minimal. It doesn't explain batch semantics (e.g., whether a single failed locus invalidates the whole call), pagination, or how the limit interacts with the batch. The deprecation note is helpful but not sufficient for a tool that likely has complex edge cases. The output schema mitigates some return-format ambiguity, but operational details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only organism has a description). The description adds semantic info for the loci parameter by noting its deprecated name (loci_or_accessions), which helps. But it provides no explanation for the limit parameter or how batching interacts with it, and the organism parameter's flexible types are not elaborated. Given the low schema coverage, more parameter guidance is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this is the batch version of string_interactions, which distinguishes it from the single-input sibling. It also specifies the batch limit (up to 50 inputs), making the tool's scope explicit. However, it doesn't restate what string_interactions does, relying on the sibling name for context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for multiple loci ('up to 50 inputs') but does not explicitly state when to choose this over the single version or how to decide between other batch tools. The deprecation note provides some usage guidance about accepted parameter names, but no exclusions or alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
biological_context_synthSynthesis: Biological ContextARead-onlyIdempotent
Synthesis: one-call equivalent of the biological_context prompt. Resolves UniProt accession, then fans out to Gramene homologs, KEGG pathways, STRING-DB partners, and ATTED-II coexpression in parallel. Adds a consensus_partners ranking that merges STRING + ATTED scores.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | ||
| top_n | No | ||
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | Synthesis tool name, e.g. analyze_locus_synth |
| input | Yes | Echoed input arguments |
| steps | Yes | Per-backend execution rows |
| result | No | Composed cross-source result; None if root step failed |
| elapsed_s | Yes | Total orchestrator wall time |
| started_at | Yes | ISO 8601 UTC timestamp |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world, and non-destructive. The description adds valuable process context beyond annotations: the resolution step, parallel fan-out, and the consensus_partners ranking merging STRING + ATTED scores. It stops short of detailing failure modes or behavior when resolution fails, but the added pipeline detail is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The core purpose and pipeline steps are front-loaded, and every clause adds distinguishing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values. It clearly defines the scope of the synthesis and its unique consensus ranking. The only notable omission is differentiation from analyze_locus_synth, another synthesis sibling whose scope could overlap; specifying the boundary would fully close the context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (organism only). The description compensates for locus by explaining it must be resolvable to a UniProt accession, and for top_n only indirectly via the consensus ranking. However, top_n semantics are still left to inference from its schema metadata (default 10, max 50), so the description does not fully bridge the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Synthesis') and a precise resource ('biological_context'), then enumerates the exact pipeline: resolves UniProt accession, fans out to Gramene, KEGG, STRING-DB, and ATTED-II, and adds a consensus_partners ranking. This clearly distinguishes it from the individual source tools among its siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Calling it a 'one-call equivalent of the biological_context prompt' and listing the four source databases implies when to use it: when a consolidated biological context is needed rather than separate calls. However, it does not explicitly state when not to use it or how it compares to sibling synthesis tools like analyze_locus_synth.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
blast_sequenceBLAST: Sequence Search (NCBI)ARead-only
Run a BLAST sequence-similarity search against NCBI BLAST URLAPI. Async Put/Get under the hood — submits the query, polls the RID (honoring NCBI's per-RID 60s floor), and returns the parsed top hits + raw text report excerpt. Programs: blastn / blastp / blastx / tblastn / tblastx. Database defaults to swissprot for protein programs, core_nt for nucleotide. Emits notifications/progress on each poll. Long searches (>10 min) raise [UpstreamUnavailableError] with the RID preserved so the client can re-poll. Set PLANT_GENOMICS_MCP_NCBI_EMAIL to identify the request per NCBI etiquette.
| Name | Required | Description | Default |
|---|---|---|---|
| expect | No | E-value threshold (default 10). | |
| program | No | BLAST program — default blastp. | blastp |
| database | No | NCBI BLAST database slug (e.g. swissprot, core_nt, refseq_protein). Defaults to swissprot for protein programs and core_nt for nucleotide programs. | |
| max_wait | No | Max seconds to wait for the search to finish before raising UpstreamUnavailableError with the RID preserved (default 600). | |
| sequence | Yes | Raw or FASTA-formatted query sequence. | |
| megablast | No | Enable megablast (blastn only). Default false. | |
| hitlist_size | No | Max hits to return (default 10). | |
| poll_interval | No | Seconds between polls. Clamped up to NCBI's per-RID 60s floor. |
Output Schema
| Name | Required | Description |
|---|---|---|
| rid | Yes | NCBI BLAST request ID — re-usable via fetch_result() |
| hits | Yes | Top alignments, sorted by BLAST default order |
| total | No | Always null: the upstream returns the top hits asked for and states no total; null means unknown, never zero (#123) |
| status | Yes | Always "READY" when this object is returned |
| program | Yes | blastn | blastp | blastx | tblastn | tblastx |
| database | Yes | NCBI BLAST database, e.g. swissprot, core_nt |
| hitCount | Yes | Number of rows parsed from the alignment summary |
| returned | Yes | Rows in this payload (#123) |
| truncated | No | Always null: without a stated total, truncation is unknown (#123) |
| elapsed_seconds | Yes | Wall-clock from submit to READY |
| raw_report_excerpt | Yes | First 50 KB of the FORMAT_TYPE=Text report |
| raw_report_truncated | Yes | True if the upstream report exceeded the cap |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give only the safety profile (readOnly, openWorld, non-destructive). The description adds substantial behavioral detail beyond them: async Put/Get polling, the NCBI per-RID 60s floor, progress notifications per poll, the >10 min failure mode raising UpstreamUnavailableError with the RID preserved for re-polling, and the required NCBI etiquette env var. This is exactly the operational context an agent needs and justifies a top score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded, opening with the core operation before moving to async mechanics, program/database defaults, and error behavior. Every sentence carries information, though the async/error/env-var block is a touch long. Efficient overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present it need not describe return values, and it doesn't over-explain them. For a long-running, polling-based external call it covers the full behavioral surface: submission, polling cadence, progress events, timeout/failure semantics with RID preservation, and NCBI identification. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters including the database default mapping and the poll floor. The description largely restates the program list and database defaults already present in the schema, adding little semantic value beyond it. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Run a BLAST sequence-similarity search against NCBI BLAST URLAPI.' It even enumerates supported programs, so the agent knows exactly what operation is performed. It stops short of differentiating itself from homology-oriented siblings (find_homologs_synth, consensus_homologs), which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer this is the tool for raw/fasta sequence input as opposed to a locus identifier. There is no explicit 'use this when you have a sequence, use X when you have a locus' routing, and no stated exclusions. Adequate but leaves selection vs siblings to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consensus_homologsSynthesis: Consensus HomologsARead-onlyIdempotent
Synthesis: cross-source homology consensus. Resolves UniProt + FASTA sequence, then runs Gramene homology calls and NCBI BLAST in parallel. Dedupes hits by normalized locus token and scores by n_sources * mean_identity — Gramene contributes identity=1.0, BLAST contributes pident/100.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | ||
| top_n | No | ||
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | Synthesis tool name, e.g. analyze_locus_synth |
| input | Yes | Echoed input arguments |
| steps | Yes | Per-backend execution rows |
| result | No | Composed cross-source result; None if root step failed |
| elapsed_s | Yes | Total orchestrator wall time |
| started_at | Yes | ISO 8601 UTC timestamp |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered by structured data. The description adds genuinely useful behavioral context beyond that: the parallel execution model, deduplication by 'normalized locus token', and the exact scoring formula including the identity=1.0 default for Gramene and pident/100 for BLAST. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Roughly 40 words, densely packed, with the purpose front-loaded in the first phrase ('cross-source homology consensus'). Every sentence earns its place: pipeline order, parallel sources, dedup rule, and scoring formula. There is zero fluff or repetition of schema/annotation content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a moderately complex multi-step pipeline, the description covers the essential mechanics: resolution step, parallel homology calls, dedup, and scoring. The output schema exists, so return-value documentation is not the description's job, and annotations cover the safety/read-only profile. Remaining gaps are the when-to-use routing (vs. find_homologs_synth) and parameter detail for locus/top_n, which were already penalized in other dimensions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% — only 'organism' has a schema description. The tool description partially compensates by explaining that the locus is 'resolved to UniProt + FASTA sequence' and that hits are scored, which implies what locus and top_n do. However, it never specifies the expected locus format (e.g., AGI ID vs symbol) or clarifies that top_n caps the returned hit list. The compensation is partial, not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs and resources: it 'resolves UniProt + FASTA sequence', 'runs Gramene homology calls and NCBI BLAST in parallel', 'dedupes hits', and 'scores by n_sources * mean_identity'. The opening phrase 'cross-source homology consensus' immediately distinguishes it from single-source siblings like gramene_homologs and blast_sequence, so an agent can tell this is the synthesis tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context — choose this when you want a consensus across multiple homology sources rather than any single one — but it never explicitly states when to use this tool versus alternatives like gramene_homologs, blast_sequence, or the similarly-named find_homologs_synth. No 'when not to use' or alternative-routing guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_plants_assemblyEnsembl Plants: AssemblyARead-onlyIdempotent
Describe an organism's Ensembl assembly (rest.ensembl.org /info/assembly; free, no key): assembly name, GCA accession and date, the karyotype, and every top-level seq-region with its length. The names are the region values ensembl_region_query takes, and a start past a region's length is refused there, so a region walk can be planned before the first call. Karyotype regions come first, in karyotype order, then unplaced scaffolds and contigs, longest first. Names are Ensembl's: tomato's chromosomes are CM001064.4 and so on, not '1'. coord_system labels differ between assemblies (chromosome, scaffold, supercontig, primary_assembly), so in_karyotype, not coord_system, says which regions are chromosomes. total counts every region before limit; truncated=true when limit cut some off (soybean has over 1,100).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| organism | Yes | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid |
Output Schema
| Name | Required | Description |
|---|---|---|
| total | Yes | How many top-level regions exist upstream for this query, all pages (pre-cap) (#123) |
| regions | Yes | Karyotype regions first, in karyotype order; then the rest, longest first |
| organism | Yes | Canonical organism slug |
| returned | Yes | Rows in this payload (#123) |
| karyotype | Yes | The chromosomes, in Ensembl's karyotype order |
| truncated | Yes | True when limit cut regions off |
| assembly_date | Yes | Assembly date as Ensembl gives it (YYYY-MM); null when Ensembl gives none |
| assembly_name | Yes | Assembly name, e.g. TAIR10, IRGSP-1.0 |
| upstream_version | No | Ensembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| assembly_accession | Yes | INSDC assembly accession (GCA_...) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, but the description goes well beyond them. It discloses ordering (karyotype first, then scaffolds/contigs longest first), naming conventions (Ensembl names like CM001064.4 vs '1'), the in_karyotype flag vs coord_system for identifying chromosomes, and the total-before-limit counting with truncated=true. These are rich behavioral traits that materially affect how an agent interprets results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences long but each sentence carries distinct weight: API/endpoint overview, integration with ensembl_region_query, ordering/naming rules, and limit/truncation behavior. It is front-loaded with the primary purpose and then layered with specifics. While not minimal, there is no filler and every clause earns its place, so a 4 fits.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is a read-only metadata endpoint with a simple 2-parameter schema and an output schema, the description covers every critical nuance an agent needs: the exact data returned, ordering rules, naming pitfalls, the distinction between coord_system and in_karyotype, and limit/truncation semantics. It even cites concrete examples (tomato, soybean) to make the behavior tangible. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema describes the organism parameter completely (slug, name, taxid) but leaves limit only with constraints. The description compensates by explaining limit's effect: 'total counts every region before limit; truncated=true when limit cut some off', and ties region lengths to downstream region_query validation. This adds practical meaning beyond the raw schema, so it earns a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Describe an organism's Ensembl assembly' and enumerates the exact output fields (assembly name, GCA accession, date, karyotype, seq-regions with lengths). It also distinguishes the tool from its sibling ensembl_region_query by explaining how the names it returns are used there, which clearly positions this tool as the assembly metadata provider.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly name alternatives or draw exclusions, but it gives a concrete usage scenario: 'a region walk can be planned before the first call' with ensembl_region_query. It implies this tool is the prerequisite for region queries and notes it is 'free, no key', which conveys straightforward access. However, it stops short of saying 'use this when X, not when Y', so it stays at 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_plants_lookup_locusEnsembl Plants: Locus MetadataARead-onlyIdempotent
Fetch metadata for a plant locus identifier from Ensembl Plants. Defaults to arabidopsis_thaliana; pass organism= for other plant species (oryza_sativa, zea_mays, ...). Locus is the TAIR-style identifier (e.g. AT1G01010 for Arabidopsis NAC001).
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | Locus identifier, e.g. AT1G01010 |
| end | No | |
| start | No | |
| source | No | |
| strand | No | 1 forward, -1 reverse |
| biotype | No | protein_coding, lncRNA, miRNA, ... |
| db_type | No | Usually "core" |
| organism | Yes | Plant organism canonical slug, e.g. arabidopsis_thaliana |
| logic_name | No | Source annotation pipeline |
| description | No | |
| object_type | No | Usually "Gene" |
| display_name | No | Human-readable gene symbol |
| assembly_name | No | e.g. TAIR10 |
| seq_region_name | No | Chromosome / contig name |
| upstream_version | No | Ensembl Plants release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl Plants did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| canonical_transcript | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the default organism and identifier format, which is useful context but not a behavioral trait beyond what annotations provide. It does not describe error behavior or edge cases, but that is not required given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences that front-load the core purpose, then add the default and identifier format. There is no redundancy or filler, and every sentence contributes to correct usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values. It covers the primary usage (fetch metadata), the default organism, and the identifier format. It doesn't mention potential limitations (e.g., which species are supported beyond the examples), but the examples and default give enough guidance for a typical call. The lack of sibling differentiation is a minor gap given the tool's simple purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for both parameters, so the baseline is 3. The description adds value by explaining the TAIR-style identifier with a concrete example (AT1G01010) and clarifying that organism defaults to arabidopsis_thaliana. This enriches the schema descriptions with practical usage context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Fetch metadata') and a specific resource ('plant locus identifier from Ensembl Plants'). It also provides an example identifier and distinguishes the data source from similar tools like TAIR or Phytozome, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the default organism and how to change it, but it never states when to prefer this tool over the many sibling tools (e.g., tair_locus_info, phytozome_lookup_locus). There is no explicit when-to-use or when-not-to-use guidance, so an agent must infer the appropriate context from the data source alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_plants_paralogsEnsembl Plants: ParalogsARead-onlyIdempotent
List the paralogues Ensembl Compara (plants) records for a plant locus (rest.ensembl.org /homology, type=paralogues; free, no key). Each row gives the paralogue's locus, its type — within_species_paralog, or other_paralog: Ensembl's 'ancient paralogues', inferred across a super tree, so the two genes can sit in different gene trees — the taxonomy_level of the duplication, perc_id/perc_pos and protein_id, closest first. gramene_homologs carries only within_species_paralog, so a gene can have paralogues here and none there. A paralogue list is not a family list: other_paralog can name a gene outside the family (AT2G23390, an acyl-CoA N-acyltransferase-like gene, is an other_paralog of the ARFs); test membership with interpro_domains. An empty list means Compara records no paralogue, not that the gene is single-copy (FLS2, AT5G46330, has none). found=false when Compara keeps no homology record for the gene at all (e.g. a non-coding gene); an id Ensembl does not know is a not-found error. total and counts_by_type count before limit; truncated=true when limit cut some off.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| locus | Yes | e.g. AT1G19850 (Arabidopsis), Os04g0519700 (rice) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | False when Compara keeps no homology record for this gene at all (e.g. a non-coding gene); true otherwise, even with no paralogue |
| locus | Yes | The locus asked about |
| total | Yes | How many paralogues exist upstream for this query, all pages (pre-cap) (#123) |
| organism | Yes | Canonical organism slug |
| paralogs | Yes | Closest first (perc_id descending) |
| returned | Yes | Rows in this payload (#123) |
| truncated | Yes | True when limit cut paralogues off |
| counts_by_type | Yes | Paralogues per type, counted before limit |
| upstream_version | No | Ensembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly, idempotent, openWorld), the description explains output row fields (locus, type, taxonomy_level, perc_id/perc_pos, protein_id), the differences between within_species_paralog and other_paralog, and the pagination behavior (total and counts_by_type before limit, truncated=true). It also warns that a paralogue list is not a family list, which is a subtle non-obvious behavior. This adds substantial context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but each sentence serves a purpose: explaining output fields, contrasting with gramene_homologs, clarifying edge cases, and describing pagination. The core purpose is front-loaded in the first sentence, and the detailed examples (AT2G23390, FLS2) justify the length. It could be slightly tighter, but the density is appropriate given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all non-obvious aspects of the tool: paralogue types, the distinction between empty list and found=false, unknown ID errors, and limit truncation. With an output schema present for return structure, the agent has everything needed to call and interpret the tool. No critical gaps remain for a wrapper around an external API.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, so locus and organism are already documented in the schema. The description does not add any semantic detail about these parameters, and only indirectly hints at the limit parameter through 'truncated=true when limit cut some off.' This is minimal added value; it could have explicitly described the limit's purpose and the organism's accepted formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'List the paralogues Ensembl Compara (plants) records for a plant locus.' It also specifies the underlying API endpoint (rest.ensembl.org /homology, type=paralogues), which unambiguously identifies the tool's function. It further differentiates from sibling gramene_homologs by noting the difference in paralogue types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description contrasts this tool with gramene_homologs, noting that a gene can have paralogues here but none there, and advises using interpro_domains to test family membership. It also clarifies critical semantics: an empty list means Compara records no paralogue, found=false means no homology record, and a not-found error for unknown IDs. This gives an agent clear guidance on when to use this tool versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_region_queryGenomic Region → FeaturesARead-onlyIdempotent
List features overlapping a genomic interval via Ensembl Plants /overlap/region. region is the seq-region name (chromosome / contig, e.g. '1'); start and end are 1-based inclusive. feature is one of gene / transcript / cds / exon (default gene). Answers 'what genes are in this QTL interval / assembly window' without a per-locus lookup. Ensembl caps the span — oversized regions error. ensembl_plants_assembly lists an organism's region names and their lengths; a region outside that list, or a start past its length, is an error here. Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | 1-based inclusive end | |
| start | Yes | 1-based start | |
| region | Yes | seq-region name (chromosome / contig), e.g. '1' or 'Chr1' | |
| feature | No | Feature type to return | gene |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | Number of overlapping features returned |
| region | Yes | seq_region:start-end, e.g. 1:3000-10000 |
| feature | Yes | Feature type queried |
| features | Yes | Raw Ensembl overlap records |
| organism | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context beyond annotations: coordinate convention, Ensembl span caps causing errors, dependency on assembly region names, and default organism behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and well-structured, front-loading the core operation, then coordinates, feature choices, use case, error behavior, and default organism in logical order. Every sentence contributes meaningful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite being a moderately complex tool, the description covers the critical context: coordinate system, feature enum, default organism, error boundaries, and relationship to the assembly tool. The presence of an output schema removes the need to describe return values, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents each parameter. The description largely restates the schema's parameter meanings with examples ('1' or 'Chr1') and the feature default, but adds little semantic detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb and resource ('List features overlapping a genomic interval via Ensembl Plants /overlap/region'), clearly stating the operation and scope. The phrase 'without a per-locus lookup' distinguishes it from sibling lookup tools, and the title reinforces the region-to-features purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes the use case: answering 'what genes are in this QTL interval / assembly window'. It also names ensembl_plants_assembly as the way to discover valid region names and lengths, and notes error conditions that help an agent decide when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entry_membersFamily / Domain MembersARead-onlyIdempotent
List every protein in one organism that carries an InterPro, Pfam or PANTHER entry, with the gene locus each maps to — the reverse of interpro_domains / panther_family (entry -> genes, not gene -> entries). One UniProt query (rest.uniprot.org; free, no key). Each member gives accession, symbol and locus; locus is the id the locus-keyed tools accept, or null when UniProt links the protein to no gene (e.g. an old cDNA submission). reviewed_only=true (default) keeps Swiss-Prot entries: complete for Arabidopsis and rice, empty for most other crops, so pass reviewed_only=false there. total is UniProt's count across all pages; when truncated=true, pass next_cursor back as cursor= for the next page. An entry absent from the organism is ok with total 0. Defaults to arabidopsis_thaliana.
| Name | Required | Description | Default |
|---|---|---|---|
| entry | Yes | InterPro (IPR010525), Pfam (PF06507) or PANTHER family (PTHR31384) | |
| cursor | No | next_cursor from the previous page; omit for the first | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
| page_size | No | ||
| reviewed_only | No | Only Swiss-Prot (curated) proteins |
Output Schema
| Name | Required | Description |
|---|---|---|
| entry | Yes | The entry asked about, e.g. IPR010525 |
| query | Yes | The UniProt query that produced this answer |
| taxid | Yes | |
| total | Yes | How many proteins matching the UniProt query exist upstream for this query, all pages (pre-cap) (#123) |
| members | Yes | |
| organism | Yes | |
| returned | Yes | Rows in this payload (#123) |
| truncated | Yes | True when more members follow next_cursor |
| next_cursor | No | Pass back as cursor= for the next page; null on the last |
| reviewed_only | Yes | |
| entry_database | Yes | |
| upstream_version | No | UniProt release that produced THIS response, as stated by its X-UniProt-Release header (e.g. '2026_02'). null means UniProt did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though annotations already mark the tool as read-only, open-world, and idempotent, the description adds substantial behavioral context: it reveals the underlying UniProt query, that it is free and keyless, what each member contains, the meaning of null locus, and the total/truncated/cursor pagination contract. This goes well beyond the annotation signals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries useful information, and the core purpose is front-loaded before the edge-case details. It reads as a compact reference that avoids filler while covering the important quirks.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description is complete: it explains the direction, required input, data source, output fields, null behavior, pagination, reviewed_only caveats, and default organism. With an output schema also present, nothing essential is missing for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers 80% of parameter meaning, and the description adds extra value by explaining reviewed_only's organism-dependent behavior and how cursor interacts with the truncated field. The description also clarifies entry types through the entry format, building on the schema examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: list every protein in an organism carrying an InterPro, Pfam, or PANTHER entry along with its gene locus. It also explicitly distinguishes itself from interpro_domains and panther_family by describing it as the reverse direction (entry -> genes).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly positions the tool against sibling tools by naming interpro_domains and panther_family and explaining the reverse relationship. It also gives concrete guidance on when to adjust reviewed_only, such as passing false for most crops because reviewed-only results are empty there.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
experimental_interactionsThaleMine: Experimental InteractionsARead-onlyIdempotent
Fetch CURATED EXPERIMENTAL protein/genetic interaction partners for an Arabidopsis locus from ThaleMine (BAR's InterMine instance; free, no key), sourced from BioGRID, IntAct and PSI-MI. Unlike string_interactions (predicted / text-mined, scored) and bar_aiv_interactions (which returns GRN paper references for Arabidopsis, not partner pairs), every partner here carries the actual experimental provenance: detection method (two hybrid, pull down, genetic interference, ...), PSI-MI relationship type, physical vs genetic class, source database, and the PubMed IDs that reported it. ThaleMine emits one row per evidence record, so rows are aggregated to one entry per partner with evidence_count as a crude support signal; partners are ordered by that count. found=false means the gene is real but has no curated interaction on record — a normal outcome; an unknown locus raises a typed NotFoundError. Arabidopsis only (ThaleMine carries genes for taxon 3702; other organisms raise OrganismNotSupported).
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | AGI locus, e.g. AT5G11260 (HY5) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid. ThaleMine supports Arabidopsis only. | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True if any curated interaction exists for this locus |
| locus | Yes | |
| total | Yes | How many interaction partners exist upstream for this query, all pages (pre-cap) (#123) |
| organism | Yes | Canonical organism slug (Arabidopsis only) |
| partners | No | Partners ordered by evidence count, descending |
| returned | Yes | Rows in this payload (#123) |
| truncated | Yes | True if the partner list was capped |
| source_url | Yes | ThaleMine gene report page — a link for the client to open or fetch; no tool on this server dereferences it |
| gene_symbol | No | Gene symbol from ThaleMine |
| partner_count | Yes | Total distinct partners (pre-cap) |
| evidence_count | Yes | Total evidence records across all partners (pre-cap) — counted over every partner upstream, not only the partners listed here |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly, openWorld, idempotent, and non-destructive behavior. The description adds substantial behavioral context: one row per evidence record is aggregated to one row per partner, partners are ordered by evidence_count, found=false semantics, NotFoundError for unknown loci, and OrganismNotSupported for non-Arabidopsis organisms. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although long, the description is dense and each sentence earns its place: source, data provenance, sibling distinctions, aggregation behavior, ordering, found=false semantics, error behavior, and organism restriction. It is front-loaded with the core purpose and then adds necessary caveats in a logical order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity and the existing annotations/output schema, the description covers all necessary operational context: data sources, result granularity, support signal, ordering, normal vs error outcomes, and organism limitations. An agent has enough information to call it correctly and interpret its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by giving an example locus (AT5G11260), clarifying that organism accepts slug/name/taxid but ThaleMine only supports Arabidopsis, and explaining how the locus parameter maps to the returned evidence. This goes beyond the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Fetch') with a precise resource: curated experimental protein/genetic interaction partners for an Arabidopsis locus. It explicitly differentiates itself from string_interactions (predicted/text-mined) and bar_aiv_interactions (GRN paper references), so an agent can distinguish it without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names alternatives and states when they are not appropriate, contrasting with string_interactions and bar_aiv_interactions. It also gives clear selection constraints: Arabidopsis only, other organisms raise OrganismNotSupported, and found=false is a normal outcome rather than an error.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
experimental_structuresPDBe: Experimental StructuresARead-onlyIdempotent
Fetch experimentally-solved (X-ray / cryo-EM / NMR) protein structures for a locus from PDBe (www.ebi.ac.uk/pdbe; free, no key). Resolves the locus → UniProt accession, then returns PDBe's best_structures mapping ranked best-first: per entry the PDB id, chain, experimental method, resolution, coverage, and modelled residue span. Most plant proteins have NO deposited structure — that returns found=false (a normal outcome, not an error); a locus with no UniProt entry raises a typed NotFoundError. structure_count is the true total even when the list is capped. Complements alphafold_structure (the predicted view). Works for all 12 organisms (UniProt-keyed). Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT4G09760 (Arabidopsis), Os01g0100100 (rice RAP-DB) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True if any experimental structure is deposited |
| locus | Yes | |
| total | Yes | How many PDBe best_structures rows (one per chain) exist upstream for this query, all pages (pre-cap) (#123) |
| returned | Yes | Rows in this payload (#123) |
| accession | Yes | Resolved UniProt accession |
| truncated | Yes | True if the structure list was capped |
| structures | No | Best-first {pdb_id, chain_id, experimental_method, resolution, coverage, …} |
| entry_count | Yes | Distinct PDB entries among the rows; structure_count counts chains (#123) |
| structure_count | Yes | PDBe rows, pre-cap — one per CHAIN, not per entry (see entry_count) |
| upstream_version | No | PDBe release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='pdbe') reports the release its own endpoint calls current at query time, or why there is none. null means PDBe did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/openWorld/idempotent/non-destructive hints, and the description adds substantial behavioral detail: locus-to-UniProt resolution, best_structures ranking and fields, found=false as a non-error, a typed NotFoundError, and the fact that structure_count is the true total even when the list is capped. This gives an agent accurate expectations for edge cases and error semantics well beyond the annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it front-loads the core purpose, then progressively covers result contents, edge cases, the relationship to alphafold_structure, organism scope, and defaults. Every clause carries useful information with no filler, though the length is at the upper edge of what is ideal for a single tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema in place, return-value details are covered elsewhere, and the description fills the remaining gaps: supported organisms, the no-structure normal case, the no-UniProt error case, the capped-list behavior, and the relationship to a sibling tool. For a two-parameter, read-only fetch with rich annotations and a full output schema, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both parameters, so the baseline is 3. The description adds meaning by explaining that results are UniProt-keyed and that locus resolution precedes the PDBe query, which clarifies how organism interacts with the lookup. It also reinforces the default organism and that organism= can be passed for other species, going slightly beyond the schema's default declaration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Fetch') and a clear resource ('experimentally-solved protein structures for a locus from PDBe'), enumerating the experimental methods (X-ray / cryo-EM / NMR). It explicitly distinguishes itself from alphafold_structure by labeling that tool as the predicted view, so there is no ambiguity about what this tool retrieves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names alphafold_structure as the predicted counterpart, implying the experimental-versus-predicted decision an agent needs to make. It also explains the expected found=false outcome for most plant proteins and calls out the NotFoundError case, which prevents an agent from misclassifying normal results as failures. The guidance is strong but the 'when-not-to-use' instruction is more implied than explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_homologs_synthSynthesis: Homolog SearchARead-onlyIdempotent
Synthesis: one-call equivalent of the find_homologs prompt. Runs BLAST then resolves UniProt-shaped subject accessions via the batch UniProt helper. Returns ranked hits each annotated with their UniProt record (or null if subject_id is not a UniProt accession).
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | ||
| program | No | blastp | |
| sequence | Yes | Query sequence (protein or nucleotide) |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | Synthesis tool name, e.g. analyze_locus_synth |
| input | Yes | Echoed input arguments |
| steps | Yes | Per-backend execution rows |
| result | No | Composed cross-source result; None if root step failed |
| elapsed_s | Yes | Total orchestrator wall time |
| started_at | Yes | ISO 8601 UTC timestamp |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds meaningful behavioral detail beyond that: it runs BLAST first, then resolves subject accessions via a batch helper, and explicitly documents the null-result case for non-UniProt accessions. This is useful and consistent with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly scoped sentences with no filler. The synthesis label and core pipeline are front-loaded, and the result semantics are stated in the final sentence, making the description easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich annotations and existing output schema, the description covers the essential pipeline, return format, and edge case (null UniProt record). It lacks explicit guidance on choosing program or interpreting top_n, but for a synthesis wrapper with this structured metadata, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low (33%), and the description does not compensate: it never explains top_n or program beyond what the schema's defaults and enum provide. The mention of 'subject_id' refers to an output concept rather than any input parameter, adding no parameter-level guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific pipeline: BLAST plus UniProt resolution, returning ranked hits with UniProt annotations. It distinguishes itself as a one-call synthesis of the find_homologs prompt and its return semantics separate it from a plain BLAST tool, though sibling differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'one-call equivalent of the find_homologs prompt' implies when to use it, and the pipeline description implies it replaces a multi-step workflow. However, it does not explicitly state when to prefer this over siblings like blast_sequence or batch_gramene_homologs, nor does it mention any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gene_reportSynthesis: Gene Report (Markdown dossier)ARead-onlyIdempotent
Synthesis: one-shot 'tell me about this gene' dossier. Resolves a locus through Ensembl Plants + UniProt, then fans out to cross-references, KEGG pathways, STRING interactors, Europe PMC literature, and QuickGO GO terms. Returns a SynthesisEnvelope whose result.markdown is a rendered Markdown gene dossier (the headline output) alongside a structured result.sections mirror; each backend payload appears once, under sections, while steps[] carries status and per-step timing only. The literature section carries no abstracts (abstracts_included: false) and the GO section no withFrom (with_from_included: false); locus_literature and locus_go_annotations return them. result.gene_names labels the Ensembl and UniProt gene names separately when they differ. Any single backend failure degrades that section to an 'Unavailable' note; the rest of the dossier still renders.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Locus name, e.g. AT1G01010 | |
| top_n | No | Caps GO terms, pathways, interactors, xrefs, and papers per section | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | Synthesis tool name, e.g. analyze_locus_synth |
| input | Yes | Echoed input arguments |
| steps | Yes | Per-backend execution rows |
| result | No | Composed cross-source result; None if root step failed |
| elapsed_s | Yes | Total orchestrator wall time |
| started_at | Yes | ISO 8601 UTC timestamp |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the (safety-only) annotations, the description discloses the output shape (result.markdown headline plus result.sections mirror, steps[] carrying only status/timing), the intentional omissions (abstracts_included: false, with_from_included: false), gene_names labeling, and graceful degradation where a single backend failure becomes an 'Unavailable' note without breaking the rest. That is exactly the behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is dense but front-loaded, leading with the one-line purpose before the fan-out and output details. Each sentence carries information, though the section/steps and omissions details are packed tightly enough that a little trimming would sharpen it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-backend synthesis tool with an output schema, the description covers the resolution pipeline, the callable surfaces, degraded-failure behavior, and the deliberate section differences versus sibling tools. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents locus, top_n (with its 1-50 cap and per-section capping semantics), and organism. The description adds no syntax or meaning beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (one-shot 'tell me about this gene' dossier) and enumerates the backend fan-out (Ensembl Plants, UniProt, KEGG, STRING, Europe PMC, QuickGO), so an agent can distinguish it from the many single-source siblings like locus_literature and locus_go_annotations without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implicitly routes the agent: it notes that locus_literature and locus_go_annotations return the abstracts and withFrom that this report omits, giving a clear condition for choosing those siblings. It lacks an explicit 'use this instead of the individual tools when…' statement or exclusions, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gene_tree_membersGene Tree MembersARead-onlyIdempotent
List the member genes of an Ensembl Compara (plants) gene tree — the gene_tree_id that gramene_homologs returns on every homolog (rest.ensembl.org /genetree; free, no key). Each member gives the locus (the id the locus-keyed tools accept), protein_id, species and taxid, and organism (the canonical slug, or null for a species outside this server's 12). target_organism= keeps one organism's members; omit it for every species. total counts members before limit; truncated=true when limit cut some off. An unknown tree id is a not-found error.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| gene_tree_id | Yes | e.g. EPlGT00940000167082 (from gramene_homologs) | |
| target_organism | No | Keep one organism's members — canonical slug, scientific or common name, or NCBI taxid. Omit for every species. |
Output Schema
| Name | Required | Description |
|---|---|---|
| total | Yes | How many members after the organism filter exist upstream for this query, all pages (pre-cap) (#123) |
| members | Yes | |
| returned | Yes | Rows in this payload (#123) |
| truncated | Yes | True when limit cut members off |
| gene_tree_id | Yes | The tree asked about, e.g. EPlGT00940000167082 |
| target_organism | Yes | The organism filter; null = every species |
| upstream_version | No | Ensembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/openWorld/idempotent annotations, the description reveals significant runtime behavior: 'free, no key', total counts members before limit, truncated=true when limit cuts results, canonical slug or null for species outside the server's 12, and 'An unknown tree id is a not-found error.' This is rich, honest behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, output contents, filtering, pagination semantics, and error behavior. It front-loads the core action and naturally orders supporting details, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with an output schema and strong annotations, the description covers everything an agent needs to call it correctly: required input provenance, optional filtering, pagination/count behavior, server scope, and failure mode. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description compensates meaningfully. It explains where gene_tree_id comes from, how target_organism can be a slug/name/taxid and that omitting it returns all species, and how limit interacts with total and truncated. All three parameters gain real semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List the member genes of an Ensembl Compara (plants) gene tree.' It clearly identifies the input (gene_tree_id from gramene_homologs) and the value of each member (locus, protein_id, species, taxid, organism), which distinguishes it from sibling tools like gramene_homologs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: this consumes the gene_tree_id that gramene_homologs returns on every homolog and produces locus ids accepted by locus-keyed tools. It does not explicitly name alternatives or state when not to use the tool, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_gene_xrefsGene Cross-ReferencesARead-onlyIdempotent
Fetch cross-database references (UniProt, NCBI Gene, TAIR, ArrayExpress, …) for a plant locus from Ensembl Plants. Defaults to arabidopsis_thaliana; pass organism= for other Ensembl Plants species. Returns count + raw xref list + a by_db rollup keyed on Ensembl's dbname (e.g. 'Uniprot_gn', 'EntrezGene') for fast lookup of a single foreign identifier.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| by_db | Yes | dbname → primary_ids[]; e.g. {'Uniprot_gn': ['Q0WV96']} |
| count | Yes | Number of xref records returned |
| locus | Yes | |
| xrefs | Yes | Raw Ensembl xref records |
| organism | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint, openWorldHint, idempotentHint) already present. Description adds return structure (count, raw xref list, by_db rollup) and dbname key explanation, going beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with action, no wasted words. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given output schema exists, description covers inputs, defaults, and output structure sufficiently. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers both parameters fully. Description adds examples (locus), explains organism accepts various forms (slug, name, taxid), enriching schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb 'Fetch' with specific resource 'cross-database references for a plant locus from Ensembl Plants'. Distinct from sibling tools like batch_get_gene_xrefs and resolve_locus_to_uniprot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides default organism and how to use for other species. Lacks explicit when-not-to-use or alternatives, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sequenceGene / CDS / Protein SequenceARead-onlyIdempotent
Fetch a locus's sequence from Ensembl Plants. seq_type is one of genomic / cds / cdna / protein (default protein — the canonical-transcript product). Closes the lookup → fetch → BLAST loop: feed the returned sequence straight to blast_sequence (protein for blastp, cds/cdna for blastn). Defaults to arabidopsis_thaliana; pass organism= for other plant species.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
| seq_type | No | Sequence type to fetch | protein |
Output Schema
| Name | Required | Description |
|---|---|---|
| type | Yes | Sequence type requested |
| locus | Yes | |
| length | Yes | Sequence length (residues for protein, bases otherwise) |
| version | No | Ensembl sequence version |
| molecule | No | "dna" or "protein" |
| organism | Yes | Resolved canonical organism slug |
| sequence | Yes | The sequence string; feed to blast_sequence |
| ensembl_id | No | Resolved Ensembl stable id |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint false. The description adds valuable behavioral context: default parameters, seq_type enum options (including that 'protein' is the canonical-transcript product), and organism flexibility. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise (two sentences of functional content plus one sentence of usage guidance) and front-loaded. Every sentence adds essential information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low parameter count, high schema coverage, and presence of output schema, the description is complete. It covers purpose, parameters with defaults, usage scenario, and chaining to sibling tools. No gaps remain for the agent to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but the description adds meaning beyond the schema by clarifying the default values for organism and seq_type, and explaining the seq_type mapping to BLAST types (protein for blastp, cds/cdna for blastn). This provides practical usage context not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches a locus's sequence from Ensembl Plants, specifying verb 'Fetch' and resource 'sequence'. It distinguishes from siblings by focusing on sequence retrieval, which is unique among the listed tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on chaining the output to blast_sequence for different seq_types, and mentions defaults for organism. It doesn't explicitly exclude alternative uses but gives strong contextual cues for when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
go_enrichmentGO / KEGG Enrichment (gene list)ARead-onlyIdempotent
GO + KEGG over-representation analysis for a gene LIST via g:Profiler g:GOSt (biit.cs.ut.ee/gprofiler; free, no API key). Unlike locus_go_annotations (one locus → its terms), this answers 'what is my gene SET enriched for?' — the dominant question for a differential-expression or co-expression cluster. loci is the query gene list (e.g. AT-codes for Arabidopsis, RAP-DB IDs for rice). sources defaults to GO:BP/GO:MF/GO:CC + KEGG; user_threshold is the g:SCS-corrected significance cutoff (default 0.05). Optional background sets a custom statistical domain (default: all annotated genes). Returns enriched[] (term_id/name/p_value/intersection_size/…, capped at top_n by p-value) plus unmapped[] — query loci g:Profiler could not recognize, surfaced so a locus-namespace mismatch is visible. Defaults to arabidopsis_thaliana; pass organism= for any of the 12 species.
| Name | Required | Description | Default |
|---|---|---|---|
| loci | Yes | Query gene set, e.g. ['AT2G46830', 'AT1G01060', ...] | |
| top_n | No | Max terms returned, sorted by p-value (1–200, default 50) | |
| sources | No | Annotation sources to test (default: all four) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
| background | No | Optional custom statistical background gene set | |
| user_threshold | No | Significance cutoff, g:SCS-corrected (default 0.05) |
Output Schema
| Name | Required | Description |
|---|---|---|
| mapped | Yes | Loci g:Profiler recognized |
| sources | Yes | Annotation sources queried |
| enriched | Yes | |
| organism | Yes | Canonical organism slug |
| returned | Yes | Terms in enriched[] after the top_n cap |
| unmapped | Yes | Loci g:Profiler could not map |
| query_size | Yes | Number of loci submitted |
| total_terms | Yes | Significant terms before the top_n cap |
| gprofiler_id | Yes | g:Profiler organism ID used, e.g. athaliana |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnlyHint, idempotentHint, etc.), the description adds significant behavioral context: it names the upstream service (g:Profiler), notes it's free and requires no API key, explains the return of unmapped loci for detecting namespace mismatches, and describes the significance correction (g:SCS). No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph with each sentence contributing essential information. It starts with the core purpose, then immediately differentiates from sibling, explains key parameters, and notes return values. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, 1 required, output schema present, extensive annotations), the description covers all essential aspects: purpose, usage context, parameter behavior, service details, return format, and organism defaults. It is fully sufficient for an AI agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions per parameter. The description adds value by explaining the role of 'loci' with examples, summarizing default behavior for sources and thresholds, and clarifying the purpose of optional parameters like background. This goes beyond the schema's structural descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs GO/KEGG over-representation analysis for a gene list, and explicitly contrasts with the sibling tool 'locus_go_annotations' which handles single loci. The verb 'answers what is my gene SET enriched for?' is specific and distinguishes the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance by contrasting with locus_go_annotations and contextualizing the tool for differential-expression or co-expression cluster analysis. It could be improved with explicit when-not-to-use scenarios, but the sibling differentiation is effective.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gramene_homologsGramene HomologsARead-onlyIdempotent
Fetch orthologs and paralogs for a plant locus from Gramene compara (data.gramene.org v69). Default homology_type='ortholog'; pass 'paralog' for in-species duplicates or 'all' for everything. 'ortholog' includes syntenic_ortholog_* rows; homoeologs (polyploid subgenome copies) come back only under 'all', and excluded_categories counts every category the filter left out. Returns target_locus + homology category (type) + shared gene_tree_id per hit. Rows carry no taxon unless with_organism=true (adds 'organism' per row) or target_organism is given, which filters to one organism before the cap and adds 'organism' per row. Paralogs here are within_species_paralog only: Gramene drops Compara's other_paralog ('ancient paralogues'); ensembl_plants_paralogs lists both. Pair with resolve_locus_to_uniprot for protein-level enrichment and with blast_sequence for sequence similarity discovery.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max homolog rows to return. 'total' always reports the true pre-cap count and 'truncated' says whether the cap bit. | |
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| cursor | No | next_cursor from the previous page; omit for the first (#123) | |
| homology_type | No | Filter on homology kind | ortholog |
| with_organism | No | Add 'organism' (Gramene species slug, null when unknown) to every row without filtering; one extra call per 100 rows (#130) | |
| target_organism | No | Keep only homologs in this organism (slug, scientific/common name, or NCBI taxid), filtered BEFORE the cap so a hub gene's rice or wheat orthologs cannot be pushed past 'limit' by other species. Adds 'organism' to every row and 'total_all_organisms'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| locus | Yes | |
| total | Yes | How many homologs (after any target_organism filter) exist upstream for this query, all pages (pre-cap) (#123) |
| release | Yes | Gramene release identifier, e.g. v69 |
| homologs | Yes | |
| returned | Yes | Rows in this payload (#123) |
| truncated | No | True when the row list was capped (< total); pass limit= to change the cap |
| next_cursor | No | Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123) |
| target_organism | No | Canonical organism the rows were filtered to, when target_organism was passed |
| upstream_version | No | Gramene release that produced THIS response, as stated by the release pinned in the request path (e.g. 'v69'); same value as release. null means Gramene did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| excluded_categories | Yes | Homologs Gramene returned that homology_type left out, counted per category (e.g. {'within_species_paralog': 3, 'homoeolog_one2one': 2} under 'ortholog'); empty under 'all'. Counted over every organism, before any target_organism filter |
| total_all_organisms | No | Homolog total before the organism filter; present only when filtered |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnly, openWorld, idempotent, and non-destructive hints; the description adds substantial behavior beyond that: homoeologs only under 'all', excluded_categories counts omitted categories, rows carry no taxon unless with_organism/target_organism is set, target_organism filters before the cap, and paralogs are within_species_paralog only. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, default behavior, edge cases, output shape, taxon behavior, tool differentiation, and integration. It is front-loaded with the core purpose and does not waffle or restate schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description is complete: it names the source version, output fields (target_locus + type + gene_tree_id), filter semantics, organism handling, cap ordering, and the key sibling alternative. The presence of an output schema reduces the need to document return details, but the description still gives a clear mental model.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes well beyond the schema by explaining non-obvious semantics: the default ortholog meaning, homoeolog behavior under 'all', target_organism's pre-cap filtering and organism-adding effect, and with_organism's per-row behavior. This materially improves parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with 'Fetch orthologs and paralogs for a plant locus from Gramene compara (data.gramene.org v69)', which names a specific verb, resource, and data source. The scope is further distinguished from sibling ensembl_plants_paralogs by stating that paralogs here are within_species_paralog only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Specifies the default homology_type='ortholog' and exactly when to pass 'paralog' or 'all'. It also names the alternative tool for the excluded case ('ensembl_plants_paralogs lists both') and recommends pairing with resolve_locus_to_uniprot and blast_sequence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
interpro_domainsInterPro: Protein DomainsARead-onlyIdempotent
Fetch the InterPro domain / family architecture for a locus (www.ebi.ac.uk/interpro; free, no key). Resolves the locus → UniProt accession, then returns the protein's InterPro entries — each with accession, name, type (domain / family / homologous_superfamily / …), source_database (Pfam appears here as source_database='pfam', not a separate tool), the integrated InterPro accession, and residue spans — plus a count_by_type rollup. A protein with no annotated domains returns found=true with an empty list; a locus with no UniProt entry raises a typed NotFoundError. domain_count is the true total even when the row list is page-capped. Works for all 12 organisms (UniProt-keyed). Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT4G09760 (Arabidopsis), Os01g0100100 (rice RAP-DB) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True once the locus resolved to a UniProt entry |
| locus | Yes | |
| total | Yes | How many InterPro entries on this protein exist upstream for this query, all pages (pre-cap) (#123) |
| domains | Yes | |
| returned | Yes | Rows in this payload (#123) |
| accession | Yes | Resolved UniProt accession |
| truncated | Yes | True if the row list was page-capped (< domain_count) |
| domain_count | Yes | Total InterPro entries (pre-cap) |
| count_by_type | Yes | Rollup of entry count by type |
| upstream_version | No | InterPro release that produced THIS response, as stated by the upstream's own header (e.g. '109.0'). null means InterPro did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Many behaviors are disclosed beyond the annotations: the locus-to-UniProt resolution workflow, the empty-list behavior for proteins without domains, the NotFoundError for loci without UniProt entries, page-capped row lists with domain_count as the true total, and the count_by_type rollup. This goes well beyond the read-only/idempotent annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds value: purpose, integration path, output shape, edge cases, pagination caveat, scope, and defaults. It is front-loaded with the core purpose, and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two well-documented parameters and an output schema, the description covers all necessary behavior: input resolution, return contents, edge cases, errors, pagination, organism scope, and default organism. Nothing an agent needs to call it correctly is left ambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents locus and organism. The description adds minor context by noting the default organism and that organism accepts other species, but it does not add substantial syntax or format guidance beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch the InterPro domain / family architecture for a locus.' It clearly enumerates what is returned (accessions, names, types, source_database, spans, count_by_type) and distinguishes itself by noting Pfam appears via source_database='pfam', not as a separate tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear selection context: use this for InterPro domain/family data, it works for all 12 organisms, and Pfam data is included here. It mentions the default organism and the organism= parameter. It does not explicitly name excluded alternatives or state when not to use this tool, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jaspar_motifJASPAR: Motif MatrixARead-onlyIdempotent
Fetch one JASPAR binding profile by matrix id, including its raw position-frequency matrix (PFM: per-base count vectors keyed A/C/G/T) plus TF class/family, assay type, source species, UniProt accessions, PubMed refs, IUPAC consensus, and the sequence-logo URL (a link for the client to fetch; no tool on this server retrieves it). The drill-down companion to tf_binding_motifs, which returns the derived consensus but not the matrix. Accepts a versioned id (MA0570.1) or a bare base id (MA0570, which resolves to the newest version). Unknown ids raise a typed NotFoundError.
| Name | Required | Description | Default |
|---|---|---|---|
| matrix_id | Yes | JASPAR profile id, e.g. MA0570.1 or MA0570 (latest version) |
Output Schema
| Name | Required | Description |
|---|---|---|
| pfm | No | Position-frequency matrix: per-base count vectors keyed A/C/G/T |
| name | No | TF name as curated by JASPAR |
| length | No | Motif width in bases |
| base_id | No | Version-less profile id, e.g. MA0570 |
| species | No | Source species [{tax_id, name}] |
| version | No | JASPAR release version of the profile |
| web_url | No | JASPAR profile page — a link for the client to open or fetch; no tool on this server dereferences it |
| tf_class | No | Structural class, e.g. ['Basic leucine zipper factors (bZIP)'] |
| consensus | No | IUPAC consensus derived from the PFM, e.g. 'AAATATCT' (the Evening Element) |
| data_type | No | Assay the profile derives from: SELEX / ChIP-seq / PBM / DAP-seq |
| matrix_id | No | JASPAR profile id, e.g. MA0570.1 |
| tf_family | No | TF family, e.g. ['MYB-related'] |
| collection | No | CORE / PBM / UNVALIDATED / … |
| pubmed_ids | No | Supporting PubMed IDs |
| uniprot_ids | No | UniProt accessions JASPAR attributes the profile to |
| sequence_logo | No | URL of the SVG sequence logo — a link for the client to open or fetch; no tool on this server dereferences it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description adds valuable behavioral details: the returned PFM structure keyed by A/C/G/T, the logo URL being a client-fetched link rather than server-retrieved, version resolution semantics, and typed NotFoundError behavior for unknown ids. These are not inferable from annotations or schema alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense but well-organized sentences front-load the primary behavior, then enumerate return contents, routing context, id semantics, and error behavior. Every clause provides selection-relevant information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with a rich output schema, the description covers the essential call semantics, id variants, return scope, an important caveat about the logo URL, and failure behavior. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the id format and versioning with 'MA0570.1 or MA0570 (latest version)'. The description repeats this idea with slightly more explicit wording about resolving to the newest version, but adds no substantively new meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Fetch one JASPAR binding profile by matrix id'. It also names the sibling tool tf_binding_motifs and distinguishes itself by returning the raw matrix, so an agent can select it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames the tool as the 'drill-down companion' to tf_binding_motifs and states the exact difference: tf_binding_motifs returns derived consensus but not the matrix. This tells the agent when this tool is the right choice versus the sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kegg_pathwaysKEGG PathwaysARead-onlyIdempotent
Fetch KEGG pathway memberships for a plant locus from rest.kegg.jp. Returns a list of pathway IDs + names + KEGG category classes the locus participates in. Pairs with locus_go_annotations for the GO-level functional view. Covers: arabidopsis_thaliana, brachypodium_distachyon, glycine_max, hordeum_vulgare, oryza_sativa, populus_trichocarpa, zea_mays. Any other organism raises OrganismNotSupported before any request (both the single and batch forms). Non-Arabidopsis loci are bridged to the NCBI Entrez Gene ID KEGG indexes (returned as entrez_gene_id). A gene KEGG knows with no pathway memberships is an ok answer with pathways=[]; a gene KEGG has no record of raises NotFoundError. KEGG v118+ is case-sensitive on the locus: pass AGI loci as uppercase.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Arabidopsis AGI locus, e.g. AT1G01010 (case preserved verbatim — KEGG v118+ is case-sensitive) | |
| organism | No | Plant organism — accepts canonical slug, scientific or common name, or NCBI taxid; see the tool description for which KEGG covers | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| locus | Yes | |
| errors | No | Per-pathway step-2 failures (kept inline so the call doesn't abort) |
| organism | Yes | Resolved canonical organism slug, e.g. arabidopsis_thaliana |
| pathways | Yes | |
| kegg_gene_id | Yes | e.g. "ath:at1g01010" |
| entrez_gene_id | No | Entrez Gene ID from the non-Arabidopsis KEGG↔Entrez bridge; absent for ath. |
| upstream_version | No | KEGG release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='kegg') reports the release its own endpoint calls current at query time, or why there is none. null means KEGG did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint, openWorldHint, idempotentHint, destructiveHint false), the description discloses several critical behaviors: OrganismNotSupported raised before any request for unsupported organisms, NotFoundError for genes KEGG has no record of, empty pathways list for genes with no memberships, case-sensitivity of KEGG v118+, and the bridging to NCBI Entrez Gene ID for non-Arabidopsis. These are valuable behavioral details that an agent needs to handle errors correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it opens with the primary action and return format, then covers supported organisms, error cases, empty-result behavior, case sensitivity, and bridging. Every sentence carries distinct information with no filler. It is longer than a minimal description but each clause earns its place; the front-loading of the core function and return type helps an agent quickly understand the purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple organisms, error conditions, case sensitivity, bridging to Entrez IDs), the description is exceptionally complete. It explains return format (pathway IDs + names + categories, plus entrez_gene_id), error semantics (OrganismNotSupported, NotFoundError, empty list), supported organisms, and the case-sensitivity requirement. With an output schema present, the agent has everything needed to call and interpret results correctly without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% coverage with descriptions for both 'locus' (case preserved verbatim, case-sensitive) and 'organism' (accepts canonical slug, scientific/common name, or NCBI taxid). The tool description adds the list of supported organisms and the non-Arabidopsis bridging note, but these are more about behavior than parameter meaning. Since schema coverage is high, the description adds marginal extra semantic value beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Fetch'), a resource ('KEGG pathway memberships') and the target ('a plant locus'), and immediately describes the return payload (pathway IDs, names, categories). It distinguishes itself from sibling tools like locus_go_annotations by mentioning it pairs with that tool, and its name clearly implies single-locus vs the batch sibling. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context by noting it pairs with locus_go_annotations for the GO-level functional view, implying when this tool is useful. However, it does not explicitly state when to use this tool vs alternatives like batch_kegg_pathways (e.g., 'use this for a single locus, use batch for multiple'), nor does it state when NOT to use it. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
locus_gene_rifsGeneRIF Functional StatementsARead-onlyIdempotent
Fetch curated GeneRIF functional statements for an Arabidopsis locus from ThaleMine (free, no key). A GeneRIF is a one-sentence, manually curated statement of what the gene does, each anchored to the PubMed ID of the publication that demonstrated it — dense, directly citable functional context that GO terms (locus_go_annotations) and raw abstracts (locus_literature) do not provide. Well-studied genes have many: HY5 (AT5G11260) has 114. Upstream order is preserved because ThaleMine supplies no meaningful ranking, so truncated means later statements were cut, not that they were less relevant. found=false means the gene exists but has no GeneRIF; an unknown locus raises a typed NotFoundError. Arabidopsis only.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | AGI locus, e.g. AT5G11260 (HY5) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid. ThaleMine supports Arabidopsis only. | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True if the gene has at least one GeneRIF |
| locus | Yes | |
| total | Yes | How many GeneRIFs exist upstream for this query, all pages (pre-cap) (#123) |
| organism | Yes | Canonical organism slug (Arabidopsis only) |
| returned | Yes | Rows in this payload (#123) |
| gene_rifs | No | Curated statements in upstream order |
| rif_count | Yes | Total GeneRIFs (pre-cap) |
| truncated | Yes | True if the GeneRIF list was capped |
| source_url | Yes | ThaleMine gene report page — a link for the client to open or fetch; no tool on this server dereferences it |
| gene_symbol | No | Gene symbol from ThaleMine |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnly/idempotent annotations already present, the description goes well beyond them: it discloses no API key required, ThaleMine provenance, preserved upstream order, the meaning of 'truncated,' the found=false semantics, and the typed NotFoundError for unknown loci. This directly informs how the agent should interpret results, and nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, and most sentences earn their place by explaining ranking, truncation, and NotFoundError behavior. It is slightly longer than strictly necessary, with minor redundancy between the opening 'Arabidopsis locus' and closing 'Arabidopsis only,' but the detail is justified by the tool's output semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with an output schema that documents the return shape, the description covers everything an agent needs to call it correctly: source, auth, organism restriction, ordering, truncation semantics, empty-result semantics, and error behavior. No significant call-time context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents both parameters thoroughly (100% coverage), including the AT5G11260 locus example and the Arabidopsis-only organism constraint, so the baseline is 3. The description adds useful context such as 'well-studied genes have many' and output edge-case semantics, but it does not materially enrich the meaning of the parameters beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch curated GeneRIF functional statements for an Arabidopsis locus from ThaleMine.' It also distinguishes itself from nearby siblings by explaining that GeneRIFs provide citable functional context that locus_go_annotations and locus_literature do not provide.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames when this tool is valuable—when dense, directly citable functional statements are needed—and names the alternatives (GO terms, raw abstracts) that fail to supply that. It also states the exclusion 'Arabidopsis only,' giving an agent a clear boundary versus multi-organism siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
locus_go_annotationsGO AnnotationsARead-onlyIdempotent
Fetch Gene Ontology annotations for a plant locus from QuickGO (EBI). Free, no API key. The locus is first resolved to a UniProt accession via the same logic as resolve_locus_to_uniprot, then QuickGO is queried by geneProductId. Returns annotations[] with goId/goName/goAspect/qualifier/evidence + a by_aspect rollup ({molecular_function: [{goId, goName}, ...], biological_process: [...], cellular_component: [...]}) deduped on goId so the high-level term set is one read away; by_aspect_deduped_on names that key in the payload. truncated is true when numberOfHits exceeds returned — raise limit (max 100).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max annotations from QuickGO (1–100, default 50) | |
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| cursor | No | next_cursor from the previous page; omit for the first (#123) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| locus | Yes | |
| total | Yes | How many GO annotations (QuickGO numberOfHits) exist upstream for this query, all pages (pre-cap) (#123) |
| returned | Yes | Rows in this payload (#123) |
| by_aspect | Yes | aspect → [{goId, goName}, ...], deduped on goId, over annotations[] only |
| truncated | Yes | True when numberOfHits exceeds returned (raise `limit`) |
| annotations | Yes | |
| next_cursor | No | Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123) |
| numberOfHits | Yes | Total annotations available upstream |
| upstream_version | No | QuickGO release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='quickgo') reports the release its own endpoint calls current at query time, or why there is none. null means QuickGO did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| uniprot_accession | Yes | UniProt accession used to query QuickGO |
| with_from_included | No | False when withFrom was not asked for (gene_report's GO section), in which case every withFrom is null because it was left out — not because the annotation has no partner or source cross-ref. |
| by_aspect_deduped_on | Yes | The key by_aspect collapses annotations[] on: a dedup, not a truncation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds real behavioral context beyond them: the UniProt resolution step, cursor paging, that 'truncated' signals numberOfHits exceeding the returned set, and that limit can be raised to 100. Error/auth behavior is the only notable omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and kept to a compact block, though the dedup/rollup sentence ('deduped on goId ... by_aspect_deduped_on names that key in the payload') is dense and partly redundant given the output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema is present, return-value explanation is optional, yet the description still covers resolution, paging, truncation, and limits, making the definition self-sufficient. Only failure/error behavior is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3), and the description meaningfully adds to it by tying 'truncated' to raising 'limit' (max 100) and clarifying the cursor is 'next_cursor from the previous page'. The organism flexibility and locus examples are already in the schema, so the extra lift is modest but real.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Fetch Gene Ontology annotations for a plant locus') plus the data source (QuickGO/EBI), and implicitly separates itself from batch_locus_go_annotations and go_enrichment by scoping to a single locus. An agent can identify what it does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the resolution pipeline and notes 'Free, no API key', which is useful context, but never states when to prefer this over batch_locus_go_annotations or go_enrichment, nor any exclusions. Usage is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
locus_literatureLiterature (Europe PMC)ARead-onlyIdempotent
Search Europe PMC for literature mentioning a plant locus. Free, no API key. Returns up to size results (default 10, capped at 25) with title, authors, journal, year, DOI, PMID, open-access status, citation count, abstract, and the article's web_url (a link for the client to open; no tool on this server retrieves it). For non-Arabidopsis species the species common name is appended to the query to disambiguate locus IDs (rice, maize, ...). Pair with resolve_locus_to_uniprot or ensembl_plants_lookup_locus to ground the locus before fanning out to the literature.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Max results (1–25, default 10) | |
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| cursor | No | next_cursor from the previous page; omit for the first (#123) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
| include_abstract | No | Set false to null out abstractText, which is ~67% of this payload. The response echoes 'abstracts_included' so a null abstract is not mistaken for an article that has none. |
Output Schema
| Name | Required | Description |
|---|---|---|
| hits | Yes | |
| locus | Yes | |
| query | Yes | Final query string sent to Europe PMC |
| total | Yes | How many papers (Europe PMC hitCount) exist upstream for this query, all pages (pre-cap) (#123) |
| hitCount | Yes | Total hits available upstream (may exceed returned) |
| organism | Yes | |
| returned | Yes | Rows in this payload (#123) |
| truncated | Yes | True when total > returned: more exist upstream than came back (#123) |
| next_cursor | No | Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123) |
| upstream_version | No | Europe PMC release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='europe_pmc') reports the release its own endpoint calls current at query time, or why there is none. null means Europe PMC did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| abstracts_included | No | False when include_abstract=False was passed, in which case every abstractText is null because it was not requested — not because the article lacks one. Abstracts are ~67% of this payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds value by stating 'Free, no API key' (cost/access), clarifying that 'web_url' is merely a link and 'no tool on this server retrieves it' (does not fetch content), and explaining the disambiguation behavior for non-Arabidopsis species. These are behavioral traits beyond the annotations, though the overall safety profile is covered by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences with no redundancy. It front-loads the core action, then lists return fields and a caveat, then explains disambiguation and pairing. Every sentence earns its place, and the structure flows logically from main purpose to details to usage context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 parameters, rich return fields) and that annotations cover safety, the description is complete. It explains the returned fields, the pagination-related 'size' behavior, the disambiguation for organism, and the absence of article retrieval via web_url. The output schema exists, so no need to repeat return structures. No critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the effect of the 'organism' parameter: 'For non-Arabidopsis species the species common name is appended to the query to disambiguate locus IDs.' It also reiterates the 'size' cap and default. This clarifies behavior not present in the parameter descriptions, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Search Europe PMC for literature mentioning a plant locus.' It clearly distinguishes from sibling tools like tair_locus_info or plantcyc_locus_info which are not about literature, and names companion tools (resolve_locus_to_uniprot, ensembl_plants_lookup_locus) for context. The verb 'search' and the resource 'Europe PMC' are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use this tool: 'Pair with resolve_locus_to_uniprot or ensembl_plants_lookup_locus to ground the locus before fanning out to the literature.' It does not explicitly mention when not to use it or alternatives like batch_locus_literature, but it implies a workflow that justifies its use. This is a clear context without explicit exclusions, fitting a score of 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
locus_plant_ontologyPlant Ontology (PO/TO) TermsARead-onlyIdempotent
Fetch Plant Ontology (PO) + Trait Ontology (TO) + experimental-condition (PECO) annotations for a plant locus from Planteome (browser.planteome.org, AmiGO2/GOlr; free, no API key). Complements locus_go_annotations: QuickGO serves GO (species-agnostic), Planteome serves the plant-specific ontologies — PO (anatomy + developmental stage), TO (traits). The locus is matched across Planteome's searchable bioentity fields and filtered by the organism's NCBI taxon. Returns annotations[] (term_id / term_name / ontology / aspect / evidence / reference) + a by_ontology rollup ({PO: [{term_id, term_name}, ...], TO: [...], PECO: [...]}) deduped on term_id. Planteome names genes by these locus ids for arabidopsis, rice, wheat and tomato only; other organisms are refused, and a gene Planteome has no record of is not found. Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max annotations from Planteome (1–200, default 100) | |
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| locus | Yes | |
| taxon | Yes | NCBI taxon filter applied, e.g. NCBITaxon:3702 |
| total | Yes | How many annotations (Planteome numFound) exist upstream for this query, all pages (pre-cap) (#123) |
| organism | Yes | Canonical organism slug |
| returned | Yes | Rows in this payload (#123) |
| truncated | Yes | True when total > returned: more exist upstream than came back (#123) |
| annotations | Yes | |
| by_ontology | Yes | namespace → [{term_id, term_name}, ...], deduped on term_id |
| numberOfHits | Yes | Total annotations available upstream |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld, non-destructive), yet the description adds substantial context beyond them: no API key required, the four-organism support limit, refusal behavior, not-found semantics, dedup-on-term_id, and the exact return shape (annotations[] plus by_ontology rollup).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and source before the sibling comparison and constraints, so nothing critical is buried. It is a single dense paragraph with several stacked clauses (return shape, dedup rule, organism limits) that could be split, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only lookup with 3 params, full schema coverage, an output schema, and rich annotations, the description covers source, auth, scope limits, defaults, and failure modes. An agent has everything needed to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning to organism beyond the schema — the default (arabidopsis_thaliana) and the fact that slug, scientific/common name, or NCBI taxid are accepted, plus which organisms are actually eligible. Limit and locus already carry adequate schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (PO/TO/PECO annotations for a plant locus) with the source (Planteome, browser.planteome.org, AmiGO2/GOlr). It explicitly distinguishes itself from the closest sibling by naming locus_go_annotations and contrasting QuickGO's species-agnostic GO with Planteome's plant-specific ontologies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent against an alternative (use locus_go_annotations for GO, this for plant-specific ontologies) and states hard conditions: only arabidopsis, rice, wheat, tomato are supported, other organisms are refused, and unknown genes are not found. Defaults and how to override organism are also given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
locus_variantsKnown VariantsARead-onlyIdempotent
List natural (germline) variants overlapping a locus's genomic span via Ensembl (rest.ensembl.org; free, no key). Resolves the locus → gene coordinates, then returns EVA/dbSNP-sourced SNPs and indels with id, source, consequence class, alleles, and clinical significance. variant_count is the true overlap total; the variant list is capped for payload size with truncated flagged. Opens the variation axis (distinct from get_sequence / ensembl_region_query). Works for all 12 organisms. Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max variant rows to return. 'variant_count' always reports the true pre-cap total and 'truncated' says whether the cap bit. | |
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice RAP-DB) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| locus | Yes | |
| total | Yes | How many variants overlapping the gene exist upstream for this query, all pages (pre-cap) (#123) |
| region | Yes | Queried gene span, e.g. '1:33666-37840' |
| gene_end | No | Gene span end (1-based) |
| organism | Yes | Resolved Ensembl species slug |
| returned | Yes | Rows in this payload (#123) |
| variants | No | Per-variant {id, source, consequence_type, alleles, …} |
| truncated | Yes | True if the variant list was capped |
| gene_start | No | Gene span start (1-based) |
| variant_count | Yes | Total overlapping variants (pre-cap) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly, openWorld, idempotent, and non-destructive behavior, and the description adds meaningful context beyond that: the variant list is capped for payload size, variant_count reports the true pre-cap total, truncated flags the cap, and the data source is EVA/dbSNP via a free keyless Ensembl endpoint. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: source, scope, resolution behavior, output fields, payload cap, sibling distinction, organism coverage, and default behavior. The most important filtering behavior is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given rich annotations, a 100%-documented input schema, and an output schema, the description covers the remaining practical concerns: data provenance, cap/truncation semantics, organism defaults, and the distinction from related region/sequence tools. Nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents locus, limit, and organism well. The description only restates the organism default and adds that all 12 organisms are supported, which is marginal value beyond the structured field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List natural (germline) variants overlapping a locus's genomic span via Ensembl'. It also names the variation axis and distinguishes the tool from get_sequence and ensembl_region_query, so an agent can differentiate it from at least its closest siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational guidance: it works for all 12 organisms, defaults to arabidopsis_thaliana, and tells the caller to pass organism= for other species. It names related tools it is 'distinct from', though it does not enumerate all when-not-to-use conditions relative to the broader sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
orthodb_orthologsOrthoDB: OrthologsARead-onlyIdempotent
Resolve a locus to its OrthoDB ortholog group and cross-species member genes (data.orthodb.org; free, no key). Searches at the Viridiplantae level, then returns the group metadata (name, evolutionary rate) and member genes grouped by organism (organism, gene id, description). organism_count is the true cluster total; the member list is capped with truncated flagged. found=false when the locus maps to no ortholog group. Works for all 12 organisms. NOTE: unlike the other locus tools, organism= does NOT scope the search — the group is resolved from the locus id alone at the Viridiplantae level, and organism is only validated and echoed back. Passing a mismatched organism therefore still returns the locus's real group. target_organism= DOES filter: it keeps only that organism's members, before the cap, so a 2,000-member group cannot hide rice or wheat behind 'limit'.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max ortholog member rows to return. 'member_count' always reports the true pre-cap total and 'truncated' says whether the cap bit. | |
| locus | Yes | e.g. AT1G01060 (Arabidopsis), Os01g0100100 (rice RAP-DB) | |
| cursor | No | next_cursor from the previous page; omit for the first (#123) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid. Validated and echoed only: it does NOT scope the OrthoDB search, which keys on the locus id at the Viridiplantae level | arabidopsis_thaliana |
| target_organism | No | Keep only this organism's members (slug, scientific/common name, or NCBI taxid), filtered BEFORE the cap. Adds 'member_count_all_organisms' for the whole group. |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True if the locus maps to an ortholog group |
| group | No | Group metadata {id, name, evolutionary_rate, level_name, …} |
| locus | Yes | |
| total | Yes | How many members (all organisms, or target_organism when given) exist upstream for this query, all pages (pre-cap) (#123) |
| members | No | Per-gene {organism, gene_id, xref, description} |
| organism | Yes | Canonical organism as requested — echoed, not inferred from the hit. Does not scope the search (see class docstring) |
| returned | Yes | Rows in this payload (#123) |
| truncated | Yes | True if the member list was capped |
| next_cursor | No | Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123) |
| member_count | Yes | Member genes before the cap (pre-cap; = total) |
| organism_count | Yes | Number of member organisms in the whole ortholog group (pre-cap) — the true cluster total, unaffected by the member cap below |
| target_organism | No | Canonical organism the members were filtered to, when target_organism was passed |
| upstream_version | No | OrthoDB release that produced THIS response, as stated by the release pinned in every request path (/v12/). null means OrthoDB did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| member_count_all_organisms | No | Whole-group member total (pre-filter, pre-cap); present only when filtered |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly, idempotent), the description adds valuable behavioral detail: the cap behavior with truncated flag, the found=false case, the exact role of organism (validated only) vs target_organism (filters before cap), and the true cluster total. This goes well beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is fairly long but every sentence adds value—covering scoping, cap behavior, and the organism distinction. It is front-loaded with the core purpose. It could be trimmed slightly, but the detail is justified given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description doesn't need to detail return values. It covers the essential behavioral nuances: the cap with truncation, the found flag, the scoping semantics, and the target_organism filtering. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents each parameter's meaning. The description largely repeats that information (e.g., the cap semantics, organism not scoping, target_organism filtering before cap). It adds little new meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Resolve'), a resource (OrthoDB ortholog group), and the result (cross-species member genes). It also distinguishes itself from sibling locus tools by noting that organism= does NOT scope the search, making its unique behavior explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains the tool's purpose (ortholog resolution) and highlights the key difference from other locus tools: 'unlike the other locus tools, organism= does NOT scope the search'. This gives agents context for when to choose this tool, though it doesn't explicitly name alternative tools or list when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
panther_familyPANTHER: Protein FamilyARead-onlyIdempotent
Fetch the PANTHER protein-family classification for a locus (pantherdb.org; free, no key). Returns the PANTHER family and subfamily (id + name) plus curated GO terms grouped by aspect (molecular_function / biological_process / cellular_component), the PANTHER protein class, and pathways. found=false when PANTHER cannot map the locus; a mapped locus can still have a null family or subfamily, which means PANTHER assigns it none. Complements the sequence-homology tools (gramene_homologs / consensus_homologs) with an evolutionary-family view. Works for all 12 organisms. Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT1G01060 (Arabidopsis), Os01g0100100 (rice RAP-DB) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True if PANTHER mapped the locus; the family fields can still be null |
| locus | Yes | |
| pathways | No | |
| accession | No | PANTHER mapped accession |
| family_id | No | PANTHER family id, e.g. PTHR12802; null when PANTHER assigns no family |
| family_name | No | |
| subfamily_id | No | e.g. PTHR12802:SF176; null when PANTHER assigns no subfamily, which is PANTHER's answer, not a failed lookup |
| protein_class | No | |
| subfamily_name | No | |
| upstream_version | No | PANTHER release that produced THIS response, as stated by the answer's search.product.version. null means PANTHER did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| go_biological_process | No | |
| go_cellular_component | No | |
| go_molecular_function | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool read-only, idempotent, and non-destructive, so the bar for added behavior is lower. The description goes well beyond annotations by disclosing failure semantics ('found=false when PANTHER cannot map the locus'), the possibility of null family/subfamily on a mapped locus, and the external source ('pantherdb.org; free, no key'). This is valuable edge-case context for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then moves logically through return contents, edge-case behavior, tool positioning, and scope/default. It is compact, about four sentences, with no filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only lookup with two well-documented parameters and an output schema, the description covers everything operationally important: what is returned, how unmapped loci behave, organism coverage and default, authentication requirements, and relation to sibling tools. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters already carry examples and defaults, so the baseline is 3. The description adds little beyond restating the organism default ('Defaults to arabidopsis_thaliana; pass organism= for other species') and provides no new syntax or format detail for either parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch the PANTHER protein-family classification for a locus', then enumerates the exact data returned (family/subfamily, GO terms, protein class, pathways). It also distinguishes itself from siblings by positioning it against gramene_homologs and consensus_homologs as an evolutionary-family view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the relevant alternatives and the relationship: 'Complements the sequence-homology tools (gramene_homologs / consensus_homologs) with an evolutionary-family view.' It also states scope and default behavior ('Works for all 12 organisms', 'Defaults to arabidopsis_thaliana'), but stops short of an explicit when-not-to-use rule or a direct 'use this instead of X' condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phytozome_lookup_locusPhytozome: Locus MetadataARead-onlyIdempotent
Fetch a gene record from Phytozome BioMart (phytozome-next.jgi.doe.gov). Defaults to arabidopsis_thaliana; pass organism= for other Phytozome proteomes (slug, scientific/common name, or NCBI taxid — e.g. glycine_max, sorghum_bicolor). Locus is the source-genome gene name (e.g. AT1G01010, Glyma.01G000100). Returns organism_name, gene_name, chromosome, gene_start, gene_end, strand, description.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Glyma.01G000100 (soybean) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| strand | Yes | String — typically "1" or "-1" |
| gene_end | Yes | String — BioMart TSV is untyped |
| gene_name | Yes | |
| chromosome | Yes | |
| gene_start | Yes | String — BioMart TSV is untyped |
| description | Yes | |
| organism_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description adds context beyond annotations (readOnlyHint, idempotentHint) by detailing the source URL and return field names. No contradictions; it complements annotation with operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences efficiently convey purpose, defaults, parameter formats, and return fields. Information is front-loaded and every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema (implied by return field list) and strong annotations, the description fully covers what the tool does, its parameters, and return values. No gaps for a lookup tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters described. Description adds examples for locus format and clarifies organism parameter accepts multiple name forms, enhancing understanding beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches a gene record from Phytozome BioMart, specifying the data source and return fields. It distinguishes from siblings like ensembl_plants_lookup_locus by naming the specific database.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Description explains defaults (arabidopsis_thaliana) and how to specify organism via slug, name, or taxid. It implies usage context for Phytozome genes but could be more explicit about when to choose this over alternative tools for other databases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plantcyc_locus_infoPlantCyc: Metabolic PathwaysARead-onlyIdempotent
Fetch metabolic annotation for a locus from PlantCyc / the Plant Metabolic Network (pmn.plantcyc.org; free BioCyc web-services API, no key). Walks gene → enzyme → catalyzed reactions → PlantCyc pathways in the organism's PGDB, returning enzymes[] + reactions[] (id/name) + pathways[] (id/name) — the metabolic-pathway view KEGG and GO don't provide. A non-enzymatic gene (e.g. a transcription factor) returns found=false with empty lists, not an error. reaction_count is the true total even when the lists are capped; pathway_count is too, or null (unknown) when the gene catalyzes more reactions than are walked for pathways. 11 organisms have a PGDB (arabidopsis, rice, maize, soybean, grape, poplar, tomato, barley, sorghum, medicago, brachypodium); wheat is not yet mapped. Defaults to arabidopsis_thaliana (AraCyc, the best-curated); pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT3G51240 (Arabidopsis), Os11g0530600 (rice RAP-DB) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True if the locus resolved to a metabolic gene |
| locus | Yes | |
| orgid | Yes | PlantCyc PGDB org id, e.g. ARA (AraCyc) |
| enzymes | Yes | Product monomer (enzyme) frame ids |
| organism | Yes | Canonical organism slug |
| pathways | Yes | |
| reactions | Yes | |
| gene_frame | No | Resolved PGDB gene frame id |
| pathway_count | Yes | Total distinct pathways (pre-cap); null (unknown) when the gene catalyzes more reactions than are walked for pathways |
| reaction_count | Yes | Total distinct reactions (pre-cap) |
| gene_common_name | No | Gene common name in the PGDB |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, open-world, idempotent, and non-destructive, so the safety profile is clear. The description adds valuable behavioral context: non-enzymatic genes return found=false with empty lists (not an error), and reaction_count/pathway_count semantics are explained (true totals even when lists are capped). This goes beyond annotations and clarifies edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it starts with the purpose, then the data flow, then edge cases, then supported organisms and defaults. Every sentence adds information; there's no fluff. It's longer than typical but justifies the length with rich, useful details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (walking gene→enzyme→reaction→pathway), the description is complete. It explains the output structure (enzymes, reactions, pathways) and return semantics (found=false for non-enzymatic genes, count meanings). It also covers organism support and default behavior. The output schema exists, so return format is documented. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters (locus and organism) with examples and description. The description adds value by explaining the default organism (arabidopsis_thaliana as AraCyc, best-curated) and that organism accepts canonical slug, scientific name, common name, or taxid. It also lists all 11 supported organisms, which is not in the schema. This goes beyond schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches metabolic annotation for a locus from PlantCyc, and walks the gene→enzyme→reaction→pathway chain to return metabolic pathways. It specifies a unique resource (PlantCyc/PMN) and distinguishes it from KEGG and GO, and the sibling list shows it's distinct from other locus tools like kegg_pathways and locus_go_annotations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool: for metabolic pathway information that KEGG and GO don't provide, and it notes that wheat is not supported (an exclusion). It doesn't explicitly say 'use this instead of X' for every sibling, but the context is clear. It also provides defaults and the organism list, which is useful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_locus_to_uniprotResolve Locus → UniProtARead-onlyIdempotent
Resolve a plant locus to its canonical UniProtKB record. Prefers reviewed (Swiss-Prot) entries; falls back to unreviewed (TrEMBL) when no curated record exists (common for non-Arabidopsis plants). organism accepts a canonical slug, scientific/common name, or NCBI taxid (default arabidopsis_thaliana; e.g. oryza_sativa, zea_mays). A gene symbol answers only when it names one locus; a symbol shared by several loci (ARF1) is InvalidArguments listing them. Returns primaryAccession, uniProtkbId, entryType, recommendedName, geneNames, organism, taxonId, sequenceLength, web_url (a link for the client to open; no tool on this server retrieves it). This is the protein-side entry point — pair with InterPro / AlphaFold / Reactome / structural-bio tools.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT1G01010 (Arabidopsis), Os01g0100100 (rice) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| taxonId | No | NCBI taxonomy ID |
| web_url | No | Browser URL for the UniProt entry — a link for the client to open or fetch; no tool on this server dereferences it |
| organism | No | Scientific name |
| reviewed | Yes | True if Swiss-Prot (curated) |
| entryType | Yes | e.g. 'UniProtKB reviewed (Swiss-Prot)' or '... (TrEMBL)' |
| geneNames | No | Gene symbols, e.g. ['NAC001'] |
| locus_query | Yes | The locus identifier the user asked about |
| uniProtkbId | Yes | UniProtKB ID, e.g. NAC1_ARATH |
| sequenceLength | No | Protein length in residues |
| recommendedName | No | Recommended protein name |
| primaryAccession | Yes | UniProt accession, e.g. Q0WV96 |
| upstream_version | No | UniProt release that produced THIS record, as stated by the upstream's own header (e.g. '2026_02'). null means UniProt did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already mark the tool as read-only, idempotent, and non-destructive. The description adds valuable behavior beyond this: reviewed-entry preference, TrEMBL fallback behavior, InvalidArguments behavior for ambiguous gene symbols, and the note that web_url is for client opening only and no server tool retrieves it. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, with the primary purpose front-loaded. Every sentence earns its place: preference behavior, parameter formats, ambiguity handling, output fields, and integration guidance are all covered without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations cover the safety profile, the description is complete. It covers input formats, default behavior, failure mode for ambiguous symbols, output highlights, and guidance on how the result link should be used. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics by explaining that locus can also be a gene symbol but only when unambiguous, and by enriching organism with defaults and concrete examples like oryza_sativa and zea_mays. This goes beyond the schema's per-parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: 'Resolve a plant locus to its canonical UniProtKB record' with explicit preference behavior for Swiss-Prot vs TrEMBL. The phrase 'protein-side entry point' further distinguishes this tool from the many locus/annotation siblings in the list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly frames this as the protein-side entry point and suggests pairing with InterPro, AlphaFold, Reactome, and structural-bio tools. It does not explicitly state when to use this over the batch sibling or the many other locus tools, but the singular/batch split and protein-side framing provide clear enough context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
string_interactionsSTRING: Interaction NetworkARead-onlyIdempotent
Fetch protein-protein interaction partners from STRING-DB (string-db.org). Accepts either a UniProt accession or a locus identifier. A locus goes to STRING's own resolver, except for wheat, whose STRING proteins carry no locus alias: a wheat locus is resolved via UniProt first. Defaults to arabidopsis_thaliana; pass organism= for other plant species (slug, scientific/common name, or NCBI taxid). Returns first-neighbor partners with the combined STRING score plus per-channel sub-scores (experimental, database, textmining, predicted). The argument was called locus_or_accession before #129; that name is still accepted, deprecated.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Number of partners to return | |
| locus | Yes | Locus (AT1G01010) or UniProt accession (Q0WV96) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| query | Yes | The locus or accession the user passed |
| total | No | Always null: the upstream returns the top partners asked for and states no total; null means unknown, never zero (#123) |
| organism | Yes | Plant organism canonical slug, e.g. arabidopsis_thaliana |
| partners | Yes | |
| returned | Yes | Rows in this payload (#123) |
| accession | Yes | UniProt accession actually queried at STRING |
| truncated | No | Always null: without a stated total, truncation is unknown (#123) |
| upstream_version | No | STRING release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='string') reports the release its own endpoint calls current at query time, or why there is none. null means STRING did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, open-world, idempotent, and non-destructive behavior. The description adds substantial behavioral detail beyond that: resolution quirks for wheat, accepted organism name forms, the returned first-neighbor scores, and the deprecated argument name that remains accepted. This gives an agent a clear mental model of observable behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is slightly long at six sentences, but every sentence adds useful information. The most important action and output are front-loaded, and the deprecated alias note is the only mildly tangential detail, though still valuable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich output schema, clear annotations, and only three parameters, the description covers all necessary invocation knowledge: accepted inputs, resolver behavior, defaults, organism selection, and return fields. Nothing essential for correct use is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema description coverage is 100%, the description enriches parameter meaning: it explains that locus may be a UniProt accession or a locus, that wheat loci are resolved via UniProt first, that organism accepts slug/scientific/common name/taxid, and that the prior argument name locus_or_accession is deprecated. This goes well beyond what the schema alone states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch protein-protein interaction partners from STRING-DB.' It further specifies accepted identifiers, default organism, and returned fields, making it easy to distinguish from sibling interaction tools such as experimental_interactions or bar_aiv_interactions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: bases of resolution (STRING resolver vs UniProt for wheat), the default organism, and how to pass other organisms. It does not explicitly name alternative tools or state when not to use this tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tair_locus_infoTAIR-Style Locus SummaryARead-onlyIdempotent
Fetch the TAIR curator-vetted Arabidopsis locus summary. Served via BAR/ThaleMine (U Toronto, Global Core Biodata Resource 2023) since TAIR's free per-locus REST API is gated behind a paid Phoenix Bioinformatics subscription. Returns TAIR curator summary + Araport11 computational description + NCBI Gene ID + cross-DB aliases (RefSeq, UniProt, TIGR locus-model IDs). Arabidopsis only. Alias of bar_gene_summary.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | Arabidopsis AGI locus, e.g. AT1G01010 |
Output Schema
| Name | Required | Description |
|---|---|---|
| agi | No | AGI primary identifier echoed by ThaleMine, e.g. "AT1G01010" |
| locus | Yes | |
| symbol | No | Gene symbol, e.g. "NAC001" |
| aliases | No | Cross-DB aliases from /gaia/aliases/ (RefSeq accessions, UniProt accessions, TIGR locus-model IDs, and TAIR aliases). Empty list if /gaia degraded. |
| species | Yes | |
| synonyms | No | TAIR aliases (CSV from Gene.tairAliases, split on commas + stripped) |
| full_name | No | Gene name from ThaleMine |
| source_url | Yes | ThaleMine endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it |
| ncbi_gene_id | No | NCBI Gene ID from /gaia/aliases/ — None if BAR has no NCBI cross-ref |
| tair_locus_id | No | TAIR locus ID from Gene.secondaryIdentifier, e.g. "locus:2200935" |
| curator_summary | No | Gene.tairCuratorSummary — the TAIR-curated functional summary prose |
| brief_description | No | Gene.briefDescription — short blurb (often same as full_name) |
| tair_short_description | No | Gene.tairShortDescription — TAIR-specific short description |
| computational_description | No | Gene.tairComputationalDescription — Araport11-sourced computed description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it is a proxy via BAR/ThaleMine, it is Arabidopsis-only, and it returns a specific set of data. It does not contradict any annotation and enriches the agent's understanding of the tool's source and constraints. Given the annotations carry the load, this is a solid 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact block of three sentences that front-loads the core purpose, then explains the serving mechanism, the return contents, and the alias. Every sentence contributes to the agent's understanding; there is no fluff. It is slightly longer than strictly necessary (the alias note could be separate), but it remains efficient and well-organized. A 4 reflects the strong structure with minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, full schema coverage, output schema present, and comprehensive annotations), the description covers the essential aspects: what it returns, its source, its scope, and its alias. It does not mention error handling or edge cases, but those are not critical for a read-only lookup. The context is complete for an agent to decide and invoke correctly, so a 4 is warranted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the single parameter 'locus' is fully described as 'Arabidopsis AGI locus, e.g. AT1G01010'. The description adds no additional parameter-level detail beyond reinforcing that it is Arabidopsis-specific, which is already implied by the schema. With full coverage, the baseline is 3, and the description doesn't need to compensate. No further elaboration is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Fetch the TAIR curator-vetted Arabidopsis locus summary.' It clearly distinguishes this tool from siblings by specifying the TAIR-specific, Arabidopsis-only scope and listing the exact content returned (curator summary, Araport11 description, NCBI Gene ID, cross-DB aliases). It also names its alias, bar_gene_summary, which further clarifies its identity. This is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the tool's purpose and why it exists (served via BAR/ThaleMine because TAIR's own API is paid), and it notes the alias relationship. This gives context for when to use it: when a TAIR-curated Arabidopsis summary is needed. However, it does not explicitly compare against other locus tools like plantcyc_locus_info or ensembl_plants_lookup_locus, nor does it state when not to use it. The guidance is clear but lacks explicit exclusions, so a 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tf_binding_motifsJASPAR: TF Binding MotifsARead-onlyIdempotent
Fetch curated transcription-factor DNA binding motifs for a locus from JASPAR (jaspar.elixir.no; free, no key) — the cis-regulatory view. Resolves the locus → UniProt accession + gene symbol, searches JASPAR by symbol scoped to the organism's taxid, then CONFIRMS each candidate by matching the accession against the profile's uniprot_ids. Returns per motif the JASPAR matrix id, TF class/family, assay type (SELEX / ChIP-seq / PBM / DAP-seq), an IUPAC consensus derived from the position-frequency matrix (e.g. CACGTG, the G-box/ABRE core), motif length, PubMed refs, and the SVG sequence_logo and JASPAR web_url — links for the client to fetch; no tool on this server retrieves them. IMPORTANT: JASPAR's name search is fuzzy, so name-similarity hits belonging to a DIFFERENT gene are returned separately in name_only_matches and must NOT be attributed to this locus; only motifs is UniProt-confirmed. found=false means the gene has no curated profile (not a TF, or its family is unprofiled for that species) — a normal outcome, not an error. Use jaspar_motif to retrieve the raw matrix for any matrix_id. Coverage is Arabidopsis-heavy (1236 profiles) and thin elsewhere (maize 131, soybean 91, wheat 58, tomato 51, rice 10; Brachypodium and sorghum have none). Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| locus | Yes | e.g. AT2G46830 (Arabidopsis CCA1), Os01g0100100 (rice RAP-DB) | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| found | Yes | True if any profile was UniProt-confirmed for this locus |
| locus | Yes | |
| total | Yes | How many UniProt-confirmed JASPAR profiles exist upstream for this query, all pages (pre-cap) (#123) |
| motifs | No | UniProt-confirmed binding profiles |
| tax_id | Yes | NCBI taxid the JASPAR search was scoped to |
| returned | Yes | Rows in this payload (#123) |
| accession | Yes | Resolved UniProt accession |
| truncated | Yes | True if the motif list was capped |
| motif_count | Yes | Total confirmed profiles (pre-cap) |
| upstream_version | No | JASPAR release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='jaspar') reports the release its own endpoint calls current at query time, or why there is none. null means JASPAR did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered. |
| name_only_matches | No | Name-similarity hits belonging to a DIFFERENT gene [{matrix_id, name, uniprot_ids}] — not this locus's motifs |
| gene_names_searched | No | Gene symbols used as JASPAR search keys |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (read-only, open world, idempotent), the description discloses the multi-step resolution process, the fuzzy-name-search caveat that returns name_only_matches separately, and the fact that the SVG/web_url are links for the client to fetch because no tool on the server retrieves them. It also explicitly labels found=false as a normal outcome, adding substantial behavioral context the annotations do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core purpose, then methodically covers process, return contents, caveats, and coverage. It is a single long paragraph but every sentence contributes; the length is justified by the tool's complexity, though slight tightening could improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters and an existing output schema, the description covers all necessary operating context: how confirmation works, what the return includes, the meaning of found=false, the fuzzy-match caveat, the server-side vs client-side fetch distinction, and species coverage. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters with 100% coverage, including examples and the default for organism. The description reinforces the default and adds coverage context but does not materially extend parameter semantics beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb+resource: 'Fetch curated transcription-factor DNA binding motifs for a locus from JASPAR' and distinguishes it from the sibling jaspar_motif by noting that this tool returns the motif collection plus metadata while jaspar_motif retrieves the raw matrix. It also identifies the cis-regulatory view, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit pointer to the alternative jaspar_motif for raw matrices and clarifies when found=false is expected (gene lacks a curated profile) rather than an error. It also provides coverage guidance across species, which helps the agent set expectations. It does not, however, enumerate conditions for preferring this over other locus tools, but in context the main alternative is addressed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upstream_releaseUpstream ReleaseARead-onlyIdempotent
Report the release a backend's own release endpoint calls current, for the backends whose answers state none (their tools' upstream_version is always null). It is read by a SEPARATE request at query time, so it is not proof of the release that answered any one call: read it before and after a run, and equal values mean no release changed in between. Never cached. ensembl_plants: the Ensembl Genomes release (e.g. '63'); string: the STRING release ('12.0'); quickgo: the GO annotation load date and the GO ontology date; jaspar: the newest ACTIVE release (several are active at once); kegg: the pathway and genes last-update dates (KEGG has no release number). pdbe, aragwas and europe_pmc publish no data release: release is null and reason says why. Backends whose answers state or pin their release (UniProt, InterPro, PANTHER, AlphaFold, Gramene, ATTED-II, OrthoDB) are not listed: read upstream_version on their answers instead.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | Yes | The backend whose current release to read |
Output Schema
| Name | Required | Description |
|---|---|---|
| reason | Yes | What the value is, or why there is none |
| backend | Yes | The backend asked about |
| release | Yes | The release its own endpoint calls current at query time — a DIFFERENT request from any data call, so not proof of the release that answered one. null when the backend publishes no data release (see reason) |
| endpoints | Yes | Full URLs of the release endpoints read, as provenance; [] when the backend publishes none — a link for the client to open or fetch; no tool on this server dereferences it |
| components | Yes | The release's parts by name, e.g. QuickGO annotation + go dates; {} when null |
| observed_at | Yes | UTC time the endpoints were read (ISO 8601) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the tool read-only, open-world, idempotent, and non-destructive. The description adds significant behavioral context beyond those hints: the release is read by a separate request at query time, it is not proof of the release that answered any one call, and it is never cached. These disclosures meaningfully prevent misuse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: it states the core purpose, flags an important query-time caveat, lists backend-specific semantics compactly, and closes with an exclusion rule. The backend mapping is structured in a readable semicolon-separated list, avoiding unnecessary prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's conceptual complexity—different meanings of 'release' per backend, null returns, and query-time semantics—the description is complete. It explains when to use this over reading upstream_version, warns about interpretation, and the output schema covers return structure, so nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter has an enum, but the description greatly enriches each enum value, explaining exactly what 'release' means per backend, e.g., STRING '12.0', quickgo's annotation and ontology dates, and kegg's last-update dates. It also clarifies that pdbe, aragwas, and europe_pmc return null with a reason because they publish no data release.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Report the release a backend's own release endpoint calls current.' It also explicitly distinguishes this tool from sibling tools by naming which backends are not listed and where to read their releases instead, so an agent can tell exactly what this tool covers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states clear when-to-use guidance: use this tool for backends whose answers state none and whose upstream_version is always null, and explicitly says when not to use it, directing the agent to read upstream_version on answers for UniProt, InterPro, PANTHER, AlphaFold, Gramene, ATTED-II, and OrthoDB. It also advises reading before and after a run to verify no release changed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vep_annotateVEP: Variant EffectARead-onlyIdempotent
Predict a variant's molecular consequences with Ensembl VEP (rest.ensembl.org; free, no key). Variant-first (not locus-first): supply an Ensembl region (chr:start-end:strand, e.g. '1:10000-10000:1') and an alternate allele (e.g. 'C'); returns the most-severe consequence plus one row per overlapping transcript (consequence terms, IMPACT, and SIFT when the variant is coding-missense; the polyphen fields stay null, as Ensembl runs PolyPhen for human only). found=false when Ensembl reports no overlapping feature. Works for all 12 organisms. Defaults to arabidopsis_thaliana; pass organism= for other species.
| Name | Required | Description | Default |
|---|---|---|---|
| allele | Yes | Alternate allele, e.g. 'C' (or 'A/C', an insertion, etc.) | |
| region | Yes | Ensembl region chr:start-end:strand, e.g. '1:10000-10000:1' | |
| organism | No | Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid | arabidopsis_thaliana |
Output Schema
| Name | Required | Description |
|---|---|---|
| end | No | |
| found | Yes | True if VEP returned an overlapping feature |
| input | No | VEP echo of the parsed input |
| start | No | |
| allele | Yes | Alternate allele, e.g. 'C' |
| region | Yes | Ensembl region, e.g. '1:10000-10000:1' |
| organism | Yes | Resolved Ensembl species slug |
| allele_string | No | |
| assembly_name | No | Assembly the call is against |
| seq_region_name | No | |
| most_severe_consequence | No | Most severe SO term |
| transcript_consequences | No | Per-transcript {gene_id, transcript_id, consequence_terms, impact, sift_*, …} |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld/non-destructive, and the description adds substantial behavior beyond them: free with no key, return shape (most-severe consequence plus one row per transcript), SIFT only when coding-missense, polyphen fields null because Ensembl runs it for human only, and found=false on no overlapping feature. This is rich, non-obvious operational detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then constraints, then output behavior, then defaults. Dense but every clause carries information; slightly long relative to a three-parameter tool but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values needn't be re-explained, yet the description still summarizes them helpfully, and it covers the organism scope, the null-polyphen caveat, and the not-found signal. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The region and allele format examples largely duplicate the schema strings; the only genuinely additive parameter context is that organism accepts slug/scientific/common name/taxid across 12 organisms and defaults to arabidopsis_thaliana. Useful but marginal over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Predict a variant's molecular consequences with Ensembl VEP') and immediately frames the scope as variant-first rather than locus-first, which is exactly the axis that separates it from the many locus_* siblings. An agent can tell what it produces (consequence terms, IMPACT, SIFT rows) without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: variant-first not locus-first, region+allele required, default organism, all 12 organisms supported. It stops short of naming a specific alternative sibling tool for the locus case, but the positive framing and the contrast clause make the when-to-use condition unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.29.0- Changed
blast_sequence1 field changed- changed
Input schema / properties / max_wait / descriptionPrevious value: -"Max seconds to wait for the search to finish before raising NotFoundError with the RID preserved (default 600)."New value: +"Max seconds to wait for the search to finish before raising UpstreamUnavailableError with the RID preserved (default 600)."
2 tool updates
v1.28.0- Changed
aragwas_associations1 field changed- added
Input schema / properties / limitAdded value: +{ + "default": 25, + "description": "Max associations, strongest first (1–100, default 25)", + "maximum": 100, + "minimum": 1, + "type": "integer" +}
- Changed
locus_go_annotations1 field changed- added
Output schema / properties / with_from_includedAdded value: +{ + "default": true, + "description": "False when withFrom was not asked for (gene_report's GO section), in which case every withFrom is null because it was left out — not because the annotation has no partner or source cross-ref.", + "title": "With From Included", + "type": "boolean" +}
14 tool updates
v1.27.0- Changed
aragwas_associations1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"AraGWAS release that produced THIS response, — always null today: this backend states no release on its responses. null means AraGWAS did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"AraGWAS release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='aragwas') reports the release its own endpoint calls current at query time, or why there is none. null means AraGWAS did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
ensembl_plants_assembly1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
ensembl_plants_lookup_locus1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"Ensembl Plants release that produced THIS response, — always null today: this backend states no release on its responses. null means Ensembl Plants did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"Ensembl Plants release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl Plants did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
ensembl_plants_paralogs1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
experimental_structures1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"PDBe release that produced THIS response, — always null today: this backend states no release on its responses. null means PDBe did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"PDBe release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='pdbe') reports the release its own endpoint calls current at query time, or why there is none. null means PDBe did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
gene_tree_members1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
kegg_pathways1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"KEGG release that produced THIS response, — always null today: this backend states no release on its responses. null means KEGG did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"KEGG release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='kegg') reports the release its own endpoint calls current at query time, or why there is none. null means KEGG did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
locus_go_annotations1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"QuickGO release that produced THIS response, — always null today: this backend states no release on its responses. null means QuickGO did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"QuickGO release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='quickgo') reports the release its own endpoint calls current at query time, or why there is none. null means QuickGO did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
locus_literature1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"Europe PMC release that produced THIS response, — always null today: this backend states no release on its responses. null means Europe PMC did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"Europe PMC release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='europe_pmc') reports the release its own endpoint calls current at query time, or why there is none. null means Europe PMC did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
orthodb_orthologs1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"OrthoDB release that produced THIS response, — always null today: this backend states no release on its responses. null means OrthoDB did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"OrthoDB release that produced THIS response, as stated by the release pinned in every request path (/v12/). null means OrthoDB did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
panther_family1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"PANTHER release that produced THIS response, — always null today: this backend states no release on its responses. null means PANTHER did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"PANTHER release that produced THIS response, as stated by the answer's search.product.version. null means PANTHER did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
string_interactions1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"STRING release that produced THIS response, — always null today: this backend states no release on its responses. null means STRING did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"STRING release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='string') reports the release its own endpoint calls current at query time, or why there is none. null means STRING did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Changed
tf_binding_motifs1 field changed- changed
Output schema / properties / upstream_version / descriptionPrevious value: -"JASPAR release that produced THIS response, — always null today: this backend states no release on its responses. null means JASPAR did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"JASPAR release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='jaspar') reports the release its own endpoint calls current at query time, or why there is none. null means JASPAR did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
- Added
upstream_release
3 tool updates
v1.26.0- Changed
batch_locus_call1 field changed- changed
Input schema / properties / tool / enumPrevious value: -[ - "alphafold_structure", - "analyze_locus_synth", - "arabidopsis_natural_variation", - "aragwas_associations", - "atted_coexpression", - "bar_aiv_interactions", - "bar_efp_expression", - "bar_gene_summary", - "biological_context_synth", - "consensus_homologs", - "ensembl_plants_lookup_locus", - "experimental_interactions", - "experimental_structures", - "gene_report", - "get_gene_xrefs", - "get_sequence", - "gramene_homologs", - "interpro_domains", - "kegg_pathways", - "locus_gene_rifs", - "locus_go_annotations", - "locus_literature", - "locus_plant_ontology", - "locus_variants", - "orthodb_orthologs", - "panther_family", - "phytozome_lookup_locus", - "plantcyc_locus_info", - "resolve_locus_to_uniprot", - "string_interactions", - "tair_locus_info", - "tf_binding_motifs" -]New value: +[ + "alphafold_structure", + "analyze_locus_synth", + "arabidopsis_natural_variation", + "aragwas_associations", + "atted_coexpression", + "bar_aiv_interactions", + "bar_efp_expression", + "bar_gene_summary", + "biological_context_synth", + "consensus_homologs", + "ensembl_plants_lookup_locus", + "ensembl_plants_paralogs", + "experimental_interactions", + "experimental_structures", + "gene_report", + "get_gene_xrefs", + "get_sequence", + "gramene_homologs", + "interpro_domains", + "kegg_pathways", + "locus_gene_rifs", + "locus_go_annotations", + "locus_literature", + "locus_plant_ontology", + "locus_variants", + "orthodb_orthologs", + "panther_family", + "phytozome_lookup_locus", + "plantcyc_locus_info", + "resolve_locus_to_uniprot", + "string_interactions", + "tair_locus_info", + "tf_binding_motifs" +]
- Added
ensembl_plants_assembly - Added
ensembl_plants_paralogs
6 tool updates
v1.24.0- Changed
atted_coexpression5 fields changed- added
Output schema / $defs / CoexNeighbor / properties / scoreAdded value: +{ + "description": "Coexpression score in the release's own index, named by the response's score_type; higher = stronger coexpression", + "title": "Score", + "type": "number" +} - changed
Output schema / $defs / CoexNeighbor / properties / z_score / descriptionPrevious value: -"ATTED-II z-score; higher = stronger coexpression"New value: +"ATTED-II z-score; higher = stronger coexpression. Null unless score_type is 'z' (Ath-u.c4-0); read score for every release" - added
Output schema / $defs / CoexNeighbor / requiredAdded value: +[ + "score" +] - added
Output schema / properties / score_typeAdded value: +{ + "description": "The coexpression index every neighbour's score is in, as ATTED-II declares it: 'z' for Ath-u.c4-0, 'LSmr' (logit score) for the other releases", + "title": "Score Type", + "type": "string" +} - changed
Output schema / requiredPrevious value: -[ - "returned", - "locus", - "atted_release", - "neighbors" -]New value: +[ + "returned", + "locus", + "atted_release", + "score_type", + "neighbors" +]
- Added
gene_tree_members - Changed
gramene_homologs3 fields changed- changed
Output schema / $defs / GrameneHomolog / properties / type / descriptionPrevious value: -"Homology category: ortholog_one2one | ortholog_one2many | ortholog_many2many | within_species_paralog | between_species_paralog"New value: +"Gramene homology category, as upstream names it, e.g. ortholog_one2one, ortholog_many2many, syntenic_ortholog_one2one, within_species_paralog, homoeolog_one2one (wheat). homology_type='ortholog' keeps the categories whose name contains 'ortholog', 'paralog' those containing 'paralog'" - added
Output schema / properties / excluded_categoriesAdded value: +{ + "additionalProperties": { + "type": "integer" + }, + "description": "Homologs Gramene returned that homology_type left out, counted per category (e.g. {'within_species_paralog': 3, 'homoeolog_one2one': 2} under 'ortholog'); empty under 'all'. Counted over every organism, before any target_organism filter", + "title": "Excluded Categories", + "type": "object" +} - changed
Output schema / requiredPrevious value: -[ - "returned", - "locus", - "release", - "total", - "homologs" -]New value: +[ + "returned", + "locus", + "release", + "total", + "homologs", + "excluded_categories" +]
- Changed
locus_literature6 fields changed- changed
Output schema / $defs / LiteratureHit / properties / hasPDF / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "type": "boolean" + }, + { + "type": "null" + } +] - changed
Output schema / $defs / LiteratureHit / properties / hasPDF / descriptionPrevious value: -"\"Y\" or \"N\""New value: +"A PDF is available (upstream Y/N)" - changed
Output schema / $defs / LiteratureHit / properties / isOpenAccess / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "type": "boolean" + }, + { + "type": "null" + } +] - changed
Output schema / $defs / LiteratureHit / properties / isOpenAccess / descriptionPrevious value: -"\"Y\" or \"N\""New value: +"Open access (upstream Y/N)" - changed
Output schema / $defs / LiteratureHit / properties / pubYear / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "type": "integer" + }, + { + "type": "null" + } +] - changed
Output schema / $defs / LiteratureHit / properties / pubYear / descriptionPrevious value: -"String — wire format is untyped"New value: +"Publication year"
- Changed
panther_family4 fields changed- changed
Output schema / descriptionPrevious value: -"PANTHER protein-family classification for a locus.\n\n``found=False`` (with null/empty fields) means PANTHER could not classify the\nlocus into a family — a normal outcome, not an error."New value: +"PANTHER protein-family classification for a locus.\n\n``found=False`` (with null/empty fields) means PANTHER could not map the\nlocus — a normal outcome, not an error. A mapped locus (``found=True``) can\nstill carry null family and subfamily fields: PANTHER knows the gene and may\nannotate GO terms for it, but assigns it no family (seen live 2026-09-25 for\nan Arabidopsis mitochondrial locus)." - changed
Output schema / properties / family_id / descriptionPrevious value: -"PANTHER family id, e.g. PTHR12802"New value: +"PANTHER family id, e.g. PTHR12802; null when PANTHER assigns no family" - changed
Output schema / properties / found / descriptionPrevious value: -"True if PANTHER classified the locus"New value: +"True if PANTHER mapped the locus; the family fields can still be null" - changed
Output schema / properties / subfamily_id / descriptionPrevious value: -"e.g. PTHR12802:SF176"New value: +"e.g. PTHR12802:SF176; null when PANTHER assigns no subfamily, which is PANTHER's answer, not a failed lookup"
- Changed
string_interactions1 field changed- removed
Output schema / $defs / StringPartner / properties / accessionRemoved value: -{ - "anyOf": [ - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Deprecated (#133): always equal to string_id, a STRING id — NOT the UniProt accession other tools mean by 'accession'. Read string_id.", - "title": "Accession" -}
38 tool updates
v1.22.0- Changed
alphafold_structure6 fields changed- changed
Output schema / properties / cif_url / descriptionPrevious value: -"mmCIF model download URL"New value: +"mmCIF model download URL — a link for the client to open or fetch; no tool on this server dereferences it" - changed
Output schema / properties / pae_image_url / descriptionPrevious value: -"Predicted-aligned-error image URL"New value: +"Predicted-aligned-error image URL — a link for the client to open or fetch; no tool on this server dereferences it" - changed
Output schema / properties / pdb_url / descriptionPrevious value: -"PDB model download URL"New value: +"PDB model download URL — a link for the client to open or fetch; no tool on this server dereferences it" - added
Output schema / properties / plddt_band_rangesAdded value: +{ + "additionalProperties": { + "items": { + "type": "integer" + }, + "type": "array" + }, + "description": "Band -> [lower, upper] pLDDT on the 0-100 scale, per EMBL-EBI: very_low <50, low 50-70, confident 70-90, very_high >90", + "title": "Plddt Band Ranges", + "type": "object" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "AlphaFold DB release that produced THIS response, as stated by the entry's own latestVersion (e.g. '6'). null means AlphaFold DB did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "accession", - "found" -]New value: +[ + "locus", + "accession", + "found", + "plddt_band_ranges" +]
- Changed
analyze_locus_synth2 fields changed- changed
Output schema / $defs / StepRow / descriptionPrevious value: -"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result``; ``status=\"error\"`` populates ``error``\nwith the existing ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)."New value: +"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result`` — unless the orchestrator carries the\npayload elsewhere in the envelope (``gene_report`` keeps it once, under\n``result.sections``), in which case the row is the audit trail alone and\n``result`` is None. ``status=\"error\"`` populates ``error`` with the\nexisting ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)." - changed
Output schema / $defs / StepRow / properties / elapsed_s / descriptionPrevious value: -"Per-step wall time when separately measurable, else None. Phase-2 gather rows and phase-0 pre-call validation failures return None because their wall time can't be honestly attributed per-step; SynthesisEnvelope.elapsed_s carries the authoritative total."New value: +"Per-step wall time: every awaited backend call is timed on its own, including inside a phase-2 gather. None only for rows that never ran (skipped, or a phase-0 pre-call validation failure); SynthesisEnvelope.elapsed_s carries the orchestrator total."
- Changed
arabidopsis_natural_variation4 fields changed- added
Input schema / properties / cursorAdded value: +{ + "description": "next_cursor from the previous page; omit for the first (#123)", + "type": "string" +} - added
Output schema / properties / next_cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123)", + "title": "Next Cursor" +} - added
Output schema / properties / totalAdded value: +{ + "description": "How many variant effects exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "found", - "transcript", - "variant_count", - "returned", - "truncated" -]New value: +[ + "total", + "locus", + "organism", + "found", + "transcript", + "variant_count", + "returned", + "truncated" +]
- Changed
aragwas_associations6 fields changed- added
Input schema / properties / cursorAdded value: +{ + "description": "next_cursor from the previous page; omit for the first (#123)", + "type": "string" +} - changed
Output schema / properties / associations / descriptionPrevious value: -"Per-hit {score, maf, mac, snp{…}, study{…}}"New value: +"Per-hit {score, maf, mac, over_bonferroni, over_fdr, over_permutation, snp{…}, study{…, thresholds}}. score is -log10(p); study.thresholds holds that study's bonferroni_threshold05/01, bh_threshold and permutation_threshold on the same scale, which over_bonferroni / over_fdr / over_permutation compare score against" - added
Output schema / properties / next_cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123)", + "title": "Next Cursor" +} - added
Output schema / properties / totalAdded value: +{ + "description": "How many associations exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "AraGWAS release that produced THIS response, — always null today: this backend states no release on its responses. null means AraGWAS did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "found", - "association_count", - "returned", - "truncated" -]New value: +[ + "total", + "locus", + "organism", + "found", + "association_count", + "returned", + "truncated" +]
- Changed
atted_coexpression5 fields changed- added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / totalAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Always null: the upstream returns the top coexpression neighbours asked for and states no total; null means unknown, never zero (#123)", + "title": "Total" +} - added
Output schema / properties / truncatedAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Always null: without a stated total, truncation is unknown (#123)", + "title": "Truncated" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "ATTED-II release that produced THIS response, as stated by the db= pinned in the request (e.g. 'Ath-u.c4-0'); same value as atted_release. null means ATTED-II did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "atted_release", - "neighbors" -]New value: +[ + "returned", + "locus", + "atted_release", + "neighbors" +]
- Changed
bar_aiv_interactions2 fields changed- changed
Output schema / $defs / BarAIVPaper / properties / image_url / descriptionPrevious value: -"BAR-hosted thumbnail of the GRN network diagram"New value: +"BAR-hosted thumbnail of the GRN network diagram — a link for the client to open or fetch; no tool on this server dereferences it" - changed
Output schema / properties / source_url / descriptionPrevious value: -"BAR AIV endpoint URL for traceability"New value: +"BAR AIV endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it"
- Changed
bar_efp_expression1 field changed- changed
Output schema / properties / source_url / descriptionPrevious value: -"BAR world-eFP endpoint URL for traceability"New value: +"BAR world-eFP endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it"
- Changed
bar_gene_summary1 field changed- changed
Output schema / properties / source_url / descriptionPrevious value: -"ThaleMine endpoint URL for traceability"New value: +"ThaleMine endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it"
- Changed
batch_atted_coexpression1 field changed- added
Input schema / properties / loci / minItemsAdded value: +1
- Changed
batch_gramene_homologs3 fields changed- added
Input schema / properties / loci / minItemsAdded value: +1 - added
Input schema / properties / target_organismAdded value: +{ + "description": "Keep only homologs in this organism (slug, scientific/common name, or NCBI taxid), filtered BEFORE the cap so a hub gene's rice or wheat orthologs cannot be pushed past 'limit' by other species. Adds 'organism' to every row and 'total_all_organisms'.", + "type": [ + "string", + "integer" + ] +} - added
Input schema / properties / with_organismAdded value: +{ + "default": false, + "description": "Add 'organism' (Gramene species slug, null when unknown) to every row without filtering; one extra call per 100 rows (#130)", + "type": "boolean" +}
- Changed
batch_kegg_pathways2 fields changed- added
Input schema / properties / loci / minItemsAdded value: +1 - changed
Input schema / properties / organism / descriptionPrevious value: -"Plant organism — only arabidopsis_thaliana is supported in v1.1.0; other plants raise OrganismNotSupported until an Entrez bridge lands"New value: +"Plant organism — accepts canonical slug, scientific or common name, or NCBI taxid; see the tool description for which KEGG covers"
- Added
batch_locus_call - Changed
batch_string_interactions3 fields changed- added
Input schema / properties / lociAdded value: +{ + "items": { + "type": "string" + }, + "maxItems": 50, + "minItems": 1, + "type": "array" +} - removed
Input schema / properties / loci_or_accessionsRemoved value: -{ - "items": { - "type": "string" - }, - "maxItems": 50, - "type": "array" -} - changed
Input schema / requiredPrevious value: -[ - "loci_or_accessions" -]New value: +[ + "loci" +]
- Changed
biological_context_synth2 fields changed- changed
Output schema / $defs / StepRow / descriptionPrevious value: -"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result``; ``status=\"error\"`` populates ``error``\nwith the existing ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)."New value: +"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result`` — unless the orchestrator carries the\npayload elsewhere in the envelope (``gene_report`` keeps it once, under\n``result.sections``), in which case the row is the audit trail alone and\n``result`` is None. ``status=\"error\"`` populates ``error`` with the\nexisting ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)." - changed
Output schema / $defs / StepRow / properties / elapsed_s / descriptionPrevious value: -"Per-step wall time when separately measurable, else None. Phase-2 gather rows and phase-0 pre-call validation failures return None because their wall time can't be honestly attributed per-step; SynthesisEnvelope.elapsed_s carries the authoritative total."New value: +"Per-step wall time: every awaited backend call is timed on its own, including inside a phase-2 gather. None only for rows that never ran (skipped, or a phase-0 pre-call validation failure); SynthesisEnvelope.elapsed_s carries the orchestrator total."
- Changed
blast_sequence4 fields changed- added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / totalAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Always null: the upstream returns the top hits asked for and states no total; null means unknown, never zero (#123)", + "title": "Total" +} - added
Output schema / properties / truncatedAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Always null: without a stated total, truncation is unknown (#123)", + "title": "Truncated" +} - changed
Output schema / requiredPrevious value: -[ - "rid", - "program", - "database", - "status", - "hitCount", - "hits", - "raw_report_excerpt", - "raw_report_truncated", - "elapsed_seconds" -]New value: +[ + "returned", + "rid", + "program", + "database", + "status", + "hitCount", + "hits", + "raw_report_excerpt", + "raw_report_truncated", + "elapsed_seconds" +]
- Changed
consensus_homologs2 fields changed- changed
Output schema / $defs / StepRow / descriptionPrevious value: -"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result``; ``status=\"error\"`` populates ``error``\nwith the existing ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)."New value: +"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result`` — unless the orchestrator carries the\npayload elsewhere in the envelope (``gene_report`` keeps it once, under\n``result.sections``), in which case the row is the audit trail alone and\n``result`` is None. ``status=\"error\"`` populates ``error`` with the\nexisting ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)." - changed
Output schema / $defs / StepRow / properties / elapsed_s / descriptionPrevious value: -"Per-step wall time when separately measurable, else None. Phase-2 gather rows and phase-0 pre-call validation failures return None because their wall time can't be honestly attributed per-step; SynthesisEnvelope.elapsed_s carries the authoritative total."New value: +"Per-step wall time: every awaited backend call is timed on its own, including inside a phase-2 gather. None only for rows that never ran (skipped, or a phase-0 pre-call validation failure); SynthesisEnvelope.elapsed_s carries the orchestrator total."
- Changed
ensembl_plants_lookup_locus1 field changed- added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Ensembl Plants release that produced THIS response, — always null today: this backend states no release on its responses. null means Ensembl Plants did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +}
- Added
entry_members - Changed
experimental_interactions4 fields changed- added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - changed
Output schema / properties / source_url / descriptionPrevious value: -"ThaleMine gene report page"New value: +"ThaleMine gene report page — a link for the client to open or fetch; no tool on this server dereferences it" - added
Output schema / properties / totalAdded value: +{ + "description": "How many interaction partners exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "found", - "partner_count", - "evidence_count", - "truncated", - "source_url" -]New value: +[ + "total", + "returned", + "locus", + "organism", + "found", + "partner_count", + "evidence_count", + "truncated", + "source_url" +]
- Changed
experimental_structures7 fields changed- changed
Output schema / descriptionPrevious value: -"PDBe experimentally-solved structures for a locus (via UniProt).\n\nThe locus is resolved to a UniProt accession, then PDBe's ``best_structures``\nmapping is queried (ranked best-first). ``found=False`` (empty list) means no\ndeposited structure — the common plant case (a 404), not an error.\nComplements ``AlphaFoldStructure`` (the predicted view). ``structure_count``\nis the true total even when the list is capped."New value: +"PDBe experimentally-solved structures for a locus (via UniProt).\n\nThe locus is resolved to a UniProt accession, then PDBe's ``best_structures``\nmapping is queried (ranked best-first). ``found=False`` (empty list) means no\ndeposited structure — the common plant case (a 404), not an error.\nComplements ``AlphaFoldStructure`` (the predicted view). ``structure_count``\n(= ``total``) counts per-chain rows before the cap; ``entry_count`` counts\ndistinct PDB entries." - added
Output schema / properties / entry_countAdded value: +{ + "description": "Distinct PDB entries among the rows; structure_count counts chains (#123)", + "title": "Entry Count", + "type": "integer" +} - added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - changed
Output schema / properties / structure_count / descriptionPrevious value: -"Total deposited structures (pre-cap)"New value: +"PDBe rows, pre-cap — one per CHAIN, not per entry (see entry_count)" - added
Output schema / properties / totalAdded value: +{ + "description": "How many PDBe best_structures rows (one per chain) exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "PDBe release that produced THIS response, — always null today: this backend states no release on its responses. null means PDBe did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "accession", - "found", - "structure_count", - "truncated" -]New value: +[ + "total", + "returned", + "entry_count", + "locus", + "accession", + "found", + "structure_count", + "truncated" +]
- Changed
find_homologs_synth2 fields changed- changed
Output schema / $defs / StepRow / descriptionPrevious value: -"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result``; ``status=\"error\"`` populates ``error``\nwith the existing ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)."New value: +"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result`` — unless the orchestrator carries the\npayload elsewhere in the envelope (``gene_report`` keeps it once, under\n``result.sections``), in which case the row is the audit trail alone and\n``result`` is None. ``status=\"error\"`` populates ``error`` with the\nexisting ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)." - changed
Output schema / $defs / StepRow / properties / elapsed_s / descriptionPrevious value: -"Per-step wall time when separately measurable, else None. Phase-2 gather rows and phase-0 pre-call validation failures return None because their wall time can't be honestly attributed per-step; SynthesisEnvelope.elapsed_s carries the authoritative total."New value: +"Per-step wall time: every awaited backend call is timed on its own, including inside a phase-2 gather. None only for rows that never ran (skipped, or a phase-0 pre-call validation failure); SynthesisEnvelope.elapsed_s carries the orchestrator total."
- Changed
gene_report2 fields changed- changed
Output schema / $defs / StepRow / descriptionPrevious value: -"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result``; ``status=\"error\"`` populates ``error``\nwith the existing ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)."New value: +"One backend call inside a synthesis envelope.\n\n``status=\"ok\"`` populates ``result`` — unless the orchestrator carries the\npayload elsewhere in the envelope (``gene_report`` keeps it once, under\n``result.sections``), in which case the row is the audit trail alone and\n``result`` is None. ``status=\"error\"`` populates ``error`` with the\nexisting ``[ExceptionClass] message`` wire format from\n``errors.PlantGenomicsError.__str__``. ``status=\"skipped\"`` populates\n``error`` with a human-readable skip reason (e.g. phase 1 failed)." - changed
Output schema / $defs / StepRow / properties / elapsed_s / descriptionPrevious value: -"Per-step wall time when separately measurable, else None. Phase-2 gather rows and phase-0 pre-call validation failures return None because their wall time can't be honestly attributed per-step; SynthesisEnvelope.elapsed_s carries the authoritative total."New value: +"Per-step wall time: every awaited backend call is timed on its own, including inside a phase-2 gather. None only for rows that never ran (skipped, or a phase-0 pre-call validation failure); SynthesisEnvelope.elapsed_s carries the orchestrator total."
- Changed
gramene_homologs11 fields changed- added
Input schema / properties / cursorAdded value: +{ + "description": "next_cursor from the previous page; omit for the first (#123)", + "type": "string" +} - added
Input schema / properties / target_organismAdded value: +{ + "description": "Keep only homologs in this organism (slug, scientific/common name, or NCBI taxid), filtered BEFORE the cap so a hub gene's rice or wheat orthologs cannot be pushed past 'limit' by other species. Adds 'organism' to every row and 'total_all_organisms'.", + "type": [ + "string", + "integer" + ] +} - added
Input schema / properties / with_organismAdded value: +{ + "default": false, + "description": "Add 'organism' (Gramene species slug, null when unknown) to every row without filtering; one extra call per 100 rows (#130)", + "type": "boolean" +} - added
Output schema / $defs / GrameneHomolog / properties / organismAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Gramene species slug of target_locus; present with with_organism or target_organism, null when Gramene has no record for the locus", + "title": "Organism" +} - added
Output schema / properties / next_cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123)", + "title": "Next Cursor" +} - added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / target_organismAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Canonical organism the rows were filtered to, when target_organism was passed", + "title": "Target Organism" +} - changed
Output schema / properties / total / descriptionPrevious value: -"Number of homologs after filtering, BEFORE the row cap"New value: +"How many homologs (after any target_organism filter) exist upstream for this query, all pages (pre-cap) (#123)" - added
Output schema / properties / total_all_organismsAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Homolog total before the organism filter; present only when filtered", + "title": "Total All Organisms" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Gramene release that produced THIS response, as stated by the release pinned in the request path (e.g. 'v69'); same value as release. null means Gramene did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "release", - "total", - "homologs" -]New value: +[ + "returned", + "locus", + "release", + "total", + "homologs" +]
- Changed
interpro_domains3 fields changed- added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / totalAdded value: +{ + "description": "How many InterPro entries on this protein exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "accession", - "found", - "domain_count", - "truncated", - "domains", - "count_by_type" -]New value: +[ + "total", + "returned", + "locus", + "accession", + "found", + "domain_count", + "truncated", + "domains", + "count_by_type" +]
- Changed
jaspar_motif2 fields changed- changed
Output schema / properties / sequence_logo / descriptionPrevious value: -"URL of the SVG sequence logo"New value: +"URL of the SVG sequence logo — a link for the client to open or fetch; no tool on this server dereferences it" - changed
Output schema / properties / web_url / descriptionPrevious value: -"JASPAR profile page"New value: +"JASPAR profile page — a link for the client to open or fetch; no tool on this server dereferences it"
- Changed
kegg_pathways2 fields changed- changed
Input schema / properties / organism / descriptionPrevious value: -"Plant organism — only arabidopsis_thaliana is supported in v1.1.0; other plants raise OrganismNotSupported until an Entrez bridge lands"New value: +"Plant organism — accepts canonical slug, scientific or common name, or NCBI taxid; see the tool description for which KEGG covers" - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "KEGG release that produced THIS response, — always null today: this backend states no release on its responses. null means KEGG did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +}
- Changed
locus_gene_rifs4 fields changed- added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - changed
Output schema / properties / source_url / descriptionPrevious value: -"ThaleMine gene report page"New value: +"ThaleMine gene report page — a link for the client to open or fetch; no tool on this server dereferences it" - added
Output schema / properties / totalAdded value: +{ + "description": "How many GeneRIFs exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "found", - "rif_count", - "truncated", - "source_url" -]New value: +[ + "total", + "returned", + "locus", + "organism", + "found", + "rif_count", + "truncated", + "source_url" +]
- Changed
locus_go_annotations9 fields changed- added
Input schema / properties / cursorAdded value: +{ + "description": "next_cursor from the previous page; omit for the first (#123)", + "type": "string" +} - changed
Output schema / properties / by_aspect / descriptionPrevious value: -"aspect → [{goId, goName}, ...], deduped on goId"New value: +"aspect → [{goId, goName}, ...], deduped on goId, over annotations[] only" - added
Output schema / properties / by_aspect_deduped_onAdded value: +{ + "const": "goId", + "description": "The key by_aspect collapses annotations[] on: a dedup, not a truncation", + "title": "By Aspect Deduped On", + "type": "string" +} - added
Output schema / properties / next_cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123)", + "title": "Next Cursor" +} - changed
Output schema / properties / returned / descriptionPrevious value: -"Number of annotations in annotations[]"New value: +"Rows in this payload (#123)" - added
Output schema / properties / totalAdded value: +{ + "description": "How many GO annotations (QuickGO numberOfHits) exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when numberOfHits exceeds returned (raise `limit`)", + "title": "Truncated", + "type": "boolean" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "QuickGO release that produced THIS response, — always null today: this backend states no release on its responses. null means QuickGO did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "uniprot_accession", - "numberOfHits", - "returned", - "annotations", - "by_aspect" -]New value: +[ + "total", + "locus", + "uniprot_accession", + "numberOfHits", + "returned", + "truncated", + "annotations", + "by_aspect", + "by_aspect_deduped_on" +]
- Changed
locus_literature8 fields changed- added
Input schema / properties / cursorAdded value: +{ + "description": "next_cursor from the previous page; omit for the first (#123)", + "type": "string" +} - changed
Output schema / $defs / LiteratureHit / properties / web_url / descriptionPrevious value: -"europepmc.org article URL"New value: +"europepmc.org article URL — a link for the client to open or fetch; no tool on this server dereferences it" - added
Output schema / properties / next_cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123)", + "title": "Next Cursor" +} - changed
Output schema / properties / returned / descriptionPrevious value: -"Number of hits actually in hits[]"New value: +"Rows in this payload (#123)" - added
Output schema / properties / totalAdded value: +{ + "description": "How many papers (Europe PMC hitCount) exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when total > returned: more exist upstream than came back (#123)", + "title": "Truncated", + "type": "boolean" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Europe PMC release that produced THIS response, — always null today: this backend states no release on its responses. null means Europe PMC did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "query", - "hitCount", - "returned", - "hits" -]New value: +[ + "total", + "truncated", + "locus", + "organism", + "query", + "hitCount", + "returned", + "hits" +]
- Changed
locus_plant_ontology4 fields changed- changed
Output schema / properties / returned / descriptionPrevious value: -"Number of annotations in annotations[]"New value: +"Rows in this payload (#123)" - added
Output schema / properties / totalAdded value: +{ + "description": "How many annotations (Planteome numFound) exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when total > returned: more exist upstream than came back (#123)", + "title": "Truncated", + "type": "boolean" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "taxon", - "numberOfHits", - "returned", - "annotations", - "by_ontology" -]New value: +[ + "total", + "truncated", + "locus", + "organism", + "taxon", + "numberOfHits", + "returned", + "annotations", + "by_ontology" +]
- Changed
locus_variants3 fields changed- added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / totalAdded value: +{ + "description": "How many variants overlapping the gene exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "region", - "variant_count", - "truncated" -]New value: +[ + "total", + "returned", + "locus", + "organism", + "region", + "variant_count", + "truncated" +]
- Changed
orthodb_orthologs10 fields changed- added
Input schema / properties / cursorAdded value: +{ + "description": "next_cursor from the previous page; omit for the first (#123)", + "type": "string" +} - added
Input schema / properties / target_organismAdded value: +{ + "description": "Keep only this organism's members (slug, scientific/common name, or NCBI taxid), filtered BEFORE the cap. Adds 'member_count_all_organisms' for the whole group.", + "type": [ + "string", + "integer" + ] +} - changed
Output schema / properties / member_count / descriptionPrevious value: -"Member genes returned (post-cap)"New value: +"Member genes before the cap (pre-cap; = total)" - added
Output schema / properties / member_count_all_organismsAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Whole-group member total (pre-filter, pre-cap); present only when filtered", + "title": "Member Count All Organisms" +} - added
Output schema / properties / next_cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pass back as cursor= to get the rows after this page; null on the last page. Opaque, and bound to this tool and query (#123)", + "title": "Next Cursor" +} - added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / target_organismAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Canonical organism the members were filtered to, when target_organism was passed", + "title": "Target Organism" +} - added
Output schema / properties / totalAdded value: +{ + "description": "How many members (all organisms, or target_organism when given) exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "OrthoDB release that produced THIS response, — always null today: this backend states no release on its responses. null means OrthoDB did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "organism", - "found", - "organism_count", - "member_count", - "truncated" -]New value: +[ + "total", + "returned", + "locus", + "organism", + "found", + "organism_count", + "member_count", + "truncated" +]
- Changed
panther_family1 field changed- added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "PANTHER release that produced THIS response, — always null today: this backend states no release on its responses. null means PANTHER did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +}
- Changed
plantcyc_locus_info4 fields changed- changed
Output schema / descriptionPrevious value: -"PlantCyc / PMN metabolic annotation for a locus.\n\nWalks gene → enzyme → reactions → pathways in the organism's PGDB via the\nfree BioCyc web-services API. ``found=False`` with empty lists when the\nlocus has no metabolic annotation (e.g. a non-enzymatic gene like a\ntranscription factor) — this is a normal result, not an error.\n``reaction_count`` / ``pathway_count`` are the true totals even when the\nreturned lists are capped (see ``plantcyc.MAX_REACTIONS`` / ``MAX_PATHWAYS``)."New value: +"PlantCyc / PMN metabolic annotation for a locus.\n\nWalks gene → enzyme → reactions → pathways in the organism's PGDB via the\nfree BioCyc web-services API. ``found=False`` with empty lists when the\nlocus has no metabolic annotation (e.g. a non-enzymatic gene like a\ntranscription factor) — this is a normal result, not an error.\n``reaction_count`` is the true total even when the returned lists are\ncapped (see ``plantcyc.MAX_REACTIONS`` / ``MAX_PATHWAYS``). ``pathway_count``\nis too, except that pathways are read from the first ``MAX_REACTIONS``\nreactions only, so past that cap it is null: unknown, never a subset's count." - added
Output schema / properties / pathway_count / anyOfAdded value: +[ + { + "type": "integer" + }, + { + "type": "null" + } +] - changed
Output schema / properties / pathway_count / descriptionPrevious value: -"Total distinct pathways (pre-cap)"New value: +"Total distinct pathways (pre-cap); null (unknown) when the gene catalyzes more reactions than are walked for pathways" - removed
Output schema / properties / pathway_count / typeRemoved value: -"integer"
- Changed
resolve_locus_to_uniprot1 field changed- changed
Output schema / properties / web_url / descriptionPrevious value: -"Browser URL for the UniProt entry"New value: +"Browser URL for the UniProt entry — a link for the client to open or fetch; no tool on this server dereferences it"
- Changed
string_interactions9 fields changed- added
Input schema / properties / locusAdded value: +{ + "description": "Locus (AT1G01010) or UniProt accession (Q0WV96)", + "type": "string" +} - removed
Input schema / properties / locus_or_accessionRemoved value: -{ - "description": "UniProt accession (Q0WV96) or locus (AT1G01010)", - "type": "string" -} - changed
Input schema / requiredPrevious value: -[ - "locus_or_accession" -]New value: +[ + "locus" +] - changed
Output schema / $defs / StringPartner / properties / accession / descriptionPrevious value: -"Partner's stringId; UniProt resolution is the caller's job"New value: +"Deprecated (#133): always equal to string_id, a STRING id — NOT the UniProt accession other tools mean by 'accession'. Read string_id." - added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / totalAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Always null: the upstream returns the top partners asked for and states no total; null means unknown, never zero (#123)", + "title": "Total" +} - added
Output schema / properties / truncatedAdded value: +{ + "anyOf": [ + { + "type": "boolean" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Always null: without a stated total, truncation is unknown (#123)", + "title": "Truncated" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "STRING release that produced THIS response, — always null today: this backend states no release on its responses. null means STRING did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "query", - "accession", - "organism", - "partners" -]New value: +[ + "returned", + "query", + "accession", + "organism", + "partners" +]
- Changed
tair_locus_info1 field changed- changed
Output schema / properties / source_url / descriptionPrevious value: -"ThaleMine endpoint URL for traceability"New value: +"ThaleMine endpoint URL for traceability — a link for the client to open or fetch; no tool on this server dereferences it"
- Changed
tf_binding_motifs6 fields changed- changed
Output schema / $defs / TfBindingMotif / properties / sequence_logo / descriptionPrevious value: -"URL of the SVG sequence logo"New value: +"URL of the SVG sequence logo — a link for the client to open or fetch; no tool on this server dereferences it" - changed
Output schema / $defs / TfBindingMotif / properties / web_url / descriptionPrevious value: -"JASPAR profile page"New value: +"JASPAR profile page — a link for the client to open or fetch; no tool on this server dereferences it" - added
Output schema / properties / returnedAdded value: +{ + "description": "Rows in this payload (#123)", + "title": "Returned", + "type": "integer" +} - added
Output schema / properties / totalAdded value: +{ + "description": "How many UniProt-confirmed JASPAR profiles exist upstream for this query, all pages (pre-cap) (#123)", + "title": "Total", + "type": "integer" +} - added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "JASPAR release that produced THIS response, — always null today: this backend states no release on its responses. null means JASPAR did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +} - changed
Output schema / requiredPrevious value: -[ - "locus", - "accession", - "tax_id", - "found", - "motif_count", - "truncated" -]New value: +[ + "total", + "returned", + "locus", + "accession", + "tax_id", + "found", + "motif_count", + "truncated" +]
2 tool updates
v1.21.0- Added
consensus_homologs - Added
gene_report
8 tool updates
v1.20.0- Removed
consensus_homologs - Removed
gene_report - Changed
gramene_homologs3 fields changed- added
Input schema / properties / limitAdded value: +{ + "default": 100, + "description": "Max homolog rows to return. 'total' always reports the true pre-cap count and 'truncated' says whether the cap bit.", + "maximum": 100, + "minimum": 1, + "type": "integer" +} - changed
Output schema / properties / total / descriptionPrevious value: -"Number of homologs after filtering"New value: +"Number of homologs after filtering, BEFORE the row cap" - added
Output schema / properties / truncatedAdded value: +{ + "default": false, + "description": "True when the row list was capped (< total); pass limit= to change the cap", + "title": "Truncated", + "type": "boolean" +}
- Changed
interpro_domains1 field changed- added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "InterPro release that produced THIS response, as stated by the upstream's own header (e.g. '109.0'). null means InterPro did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +}
- Changed
locus_literature2 fields changed- added
Input schema / properties / include_abstractAdded value: +{ + "default": true, + "description": "Set false to null out abstractText, which is ~67% of this payload. The response echoes 'abstracts_included' so a null abstract is not mistaken for an article that has none.", + "type": "boolean" +} - added
Output schema / properties / abstracts_includedAdded value: +{ + "default": true, + "description": "False when include_abstract=False was passed, in which case every abstractText is null because it was not requested — not because the article lacks one. Abstracts are ~67% of this payload.", + "title": "Abstracts Included", + "type": "boolean" +}
- Changed
locus_variants1 field changed- added
Input schema / properties / limitAdded value: +{ + "default": 500, + "description": "Max variant rows to return. 'variant_count' always reports the true pre-cap total and 'truncated' says whether the cap bit.", + "maximum": 500, + "minimum": 1, + "type": "integer" +}
- Changed
orthodb_orthologs1 field changed- added
Input schema / properties / limitAdded value: +{ + "default": 100, + "description": "Max ortholog member rows to return. 'member_count' always reports the true pre-cap total and 'truncated' says whether the cap bit.", + "maximum": 100, + "minimum": 1, + "type": "integer" +}
- Changed
resolve_locus_to_uniprot1 field changed- added
Output schema / properties / upstream_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "UniProt release that produced THIS record, as stated by the upstream's own header (e.g. '2026_02'). null means UniProt did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.", + "title": "Upstream Version" +}
18 tool updates
v1.19.4- Added
aragwas_associations - Added
atted_coexpression - Changed
batch_atted_coexpression1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_bar_aiv_interactions1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_bar_gene_summary1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_ensembl_plants_lookup_locus1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Added
batch_get_gene_xrefs - Changed
batch_gramene_homologs1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_kegg_pathways1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_locus_go_annotations1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_locus_literature1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_phytozome_lookup_locus1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_resolve_locus_to_uniprot1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
batch_string_interactions1 field changed- changed
Output schema / properties / count / descriptionPrevious value: -"Number of loci in the input list"New value: +"Number of distinct loci queried, returned (== len(results) + len(errors)). The input list is de-duplicated first, so this is LOWER than the number of loci you sent if you sent a duplicate — that is de-duplication, not a dropped locus."
- Changed
experimental_interactions1 field changed- changed
Output schema / properties / evidence_count / descriptionPrevious value: -"Total evidence records across all partners"New value: +"Total evidence records across all partners (pre-cap) — counted over every partner upstream, not only the partners listed here"
- Added
locus_literature - Added
locus_plant_ontology - Changed
orthodb_orthologs4 fields changed- changed
Input schema / properties / organism / descriptionPrevious value: -"Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid"New value: +"Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid. Validated and echoed only: it does NOT scope the OrthoDB search, which keys on the locus id at the Viridiplantae level" - changed
Output schema / descriptionPrevious value: -"OrthoDB ortholog group + cross-species member genes for a locus.\n\n``found=False`` means the locus maps to no Viridiplantae ortholog group.\n``organism_count`` is the true cluster total even when members are capped."New value: +"OrthoDB ortholog group + cross-species member genes for a locus.\n\n``found=False`` means the locus maps to no Viridiplantae ortholog group.\n``organism_count`` is the true cluster total even when members are capped.\n\n``organism`` is an echo of the request, NOT a property of the result: the\nsearch keys on the locus id at the Viridiplantae level, so the group comes\nback the same whichever organism was declared." - changed
Output schema / properties / organism / descriptionPrevious value: -"Resolved canonical organism"New value: +"Canonical organism as requested — echoed, not inferred from the hit. Does not scope the search (see class docstring)" - changed
Output schema / properties / organism_count / descriptionPrevious value: -"Number of member organisms (clusters)"New value: +"Number of member organisms in the whole ortholog group (pre-cap) — the true cluster total, unaffected by the member cap below"
20 tool updates
v1.18.2- Added
alphafold_structure - Added
arabidopsis_natural_variation - Removed
atted_coexpression - Removed
batch_get_gene_xrefs - Added
ensembl_region_query - Added
experimental_interactions - Added
experimental_structures - Added
gene_report - Added
get_sequence - Added
go_enrichment - Added
interpro_domains - Added
jaspar_motif - Added
locus_gene_rifs - Removed
locus_literature - Added
locus_variants - Added
orthodb_orthologs - Added
panther_family - Changed
plantcyc_locus_info21 fields changed- changed
Input schema / properties / locus / descriptionPrevious value: -"TAIR-canonical locus, e.g. AT1G01010"New value: +"e.g. AT3G51240 (Arabidopsis), Os11g0530600 (rice RAP-DB)" - added
Input schema / properties / organismAdded value: +{ + "default": "arabidopsis_thaliana", + "description": "Plant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid", + "type": [ + "string", + "integer" + ] +} - added
Output schema / $defsAdded value: +{ + "PlantCycPathway": { + "additionalProperties": true, + "description": "One PlantCyc/PMN pathway the locus participates in.", + "properties": { + "id": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pathway frame id, e.g. PWY-6787", + "title": "Id" + }, + "name": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Pathway common name, e.g. flavonoid biosynthesis", + "title": "Name" + } + }, + "title": "PlantCycPathway", + "type": "object" + }, + "PlantCycReaction": { + "additionalProperties": true, + "description": "One reaction catalyzed by a locus's gene product (PlantCyc/PMN).", + "properties": { + "id": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Reaction frame id, e.g. RXN-7775", + "title": "Id" + }, + "name": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Reaction common name, if the frame has one", + "title": "Name" + } + }, + "title": "PlantCycReaction", + "type": "object" + } +} - changed
Output schema / descriptionPrevious value: -"PlantCyc stub response — adds ``plantcyc_web_url`` to the shared shape."New value: +"PlantCyc / PMN metabolic annotation for a locus.\n\nWalks gene → enzyme → reactions → pathways in the organism's PGDB via the\nfree BioCyc web-services API. ``found=False`` with empty lists when the\nlocus has no metabolic annotation (e.g. a non-enzymatic gene like a\ntranscription factor) — this is a normal result, not an error.\n``reaction_count`` / ``pathway_count`` are the true totals even when the\nreturned lists are capped (see ``plantcyc.MAX_REACTIONS`` / ``MAX_PATHWAYS``)." - removed
Output schema / properties / alternativesRemoved value: -{ - "description": "Tool names users should call instead", - "items": { - "type": "string" - }, - "title": "Alternatives", - "type": "array" -} - removed
Output schema / properties / alternatives_noteRemoved value: -{ - "description": "What the alternatives do and do NOT cover", - "title": "Alternatives Note", - "type": "string" -} - added
Output schema / properties / enzymesAdded value: +{ + "description": "Product monomer (enzyme) frame ids", + "items": { + "type": "string" + }, + "title": "Enzymes", + "type": "array" +} - added
Output schema / properties / foundAdded value: +{ + "description": "True if the locus resolved to a metabolic gene", + "title": "Found", + "type": "boolean" +} - added
Output schema / properties / gene_common_nameAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Gene common name in the PGDB", + "title": "Gene Common Name" +} - added
Output schema / properties / gene_frameAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Resolved PGDB gene frame id", + "title": "Gene Frame" +} - added
Output schema / properties / organismAdded value: +{ + "description": "Canonical organism slug", + "title": "Organism", + "type": "string" +} - added
Output schema / properties / orgidAdded value: +{ + "description": "PlantCyc PGDB org id, e.g. ARA (AraCyc)", + "title": "Orgid", + "type": "string" +} - added
Output schema / properties / pathway_countAdded value: +{ + "description": "Total distinct pathways (pre-cap)", + "title": "Pathway Count", + "type": "integer" +} - added
Output schema / properties / pathwaysAdded value: +{ + "items": { + "$ref": "#/$defs/PlantCycPathway" + }, + "title": "Pathways", + "type": "array" +} - removed
Output schema / properties / plantcyc_web_urlRemoved value: -{ - "description": "Browser URL for the PlantCyc gene page", - "title": "Plantcyc Web Url", - "type": "string" -} - removed
Output schema / properties / probed_atRemoved value: -{ - "description": "ISO date of the last live access probe (YYYY-MM-DD)", - "title": "Probed At", - "type": "string" -} - removed
Output schema / properties / rationaleRemoved value: -{ - "description": "Why this backend is gated", - "title": "Rationale", - "type": "string" -} - added
Output schema / properties / reaction_countAdded value: +{ + "description": "Total distinct reactions (pre-cap)", + "title": "Reaction Count", + "type": "integer" +} - added
Output schema / properties / reactionsAdded value: +{ + "items": { + "$ref": "#/$defs/PlantCycReaction" + }, + "title": "Reactions", + "type": "array" +} - removed
Output schema / properties / statusRemoved value: -{ - "description": "Always \"subscription_required\" — upstream REST is paid-only.", - "title": "Status", - "type": "string" -} - changed
Output schema / requiredPrevious value: -[ - "locus", - "status", - "probed_at", - "rationale", - "alternatives", - "alternatives_note", - "plantcyc_web_url" -]New value: +[ + "locus", + "organism", + "orgid", + "found", + "enzymes", + "reactions", + "pathways", + "reaction_count", + "pathway_count" +]
- Added
tf_binding_motifs - Added
vep_annotate
32 tool updates
v1.8.0- First observed
analyze_locus_synth - First observed
atted_coexpression - First observed
bar_aiv_interactions - First observed
bar_efp_expression - First observed
bar_gene_summary - First observed
batch_atted_coexpression - First observed
batch_bar_aiv_interactions - First observed
batch_bar_gene_summary - First observed
batch_ensembl_plants_lookup_locus - First observed
batch_get_gene_xrefs - First observed
batch_gramene_homologs - First observed
batch_kegg_pathways - First observed
batch_locus_go_annotations - First observed
batch_locus_literature - First observed
batch_phytozome_lookup_locus - First observed
batch_resolve_locus_to_uniprot - First observed
batch_string_interactions - First observed
biological_context_synth - First observed
blast_sequence - First observed
consensus_homologs - First observed
ensembl_plants_lookup_locus - First observed
find_homologs_synth - First observed
get_gene_xrefs - First observed
gramene_homologs - First observed
kegg_pathways - First observed
locus_go_annotations - First observed
locus_literature - First observed
phytozome_lookup_locus - First observed
plantcyc_locus_info - First observed
resolve_locus_to_uniprot - First observed
string_interactions - First observed
tair_locus_info
TDQS
Scored across 56 tools
The descriptions are unusually thorough and often explicitly contrast sibling tools (e.g., string_interactions vs experimental_interactions, alphafold_structure vs experimental_structures). However, the set contains several genuinely overlapping clusters: the tair_locus_info/bar_gene_summary alias, multiple synthesis dossiers (gene_report, analyze_locus_synth, biological_context_synth, consensus_homologs), many homology/orthology tools, and the generic batch_locus_call alongside dedicated batch_* variants. An agent could still misselect among these, especially when choosing between a one-shot dossier and its component tools.
Most tools use snake_case with recognizable patterns: batch_* for batch variants, resource_action or locus_attribute for single-source tools, and _synth for prompt-equivalent syntheses. Minor deviations exist, such as verb-first names (get_sequence, resolve_locus_to_uniprot, find_homologs_synth) mixed with noun-first names (locus_go_annotations, interpro_domains, blast_sequence), but the convention remains readable and largely predictable. This is not chaotic, just slightly mixed.
56 tools is far beyond the typical 3-15 range and is inflated by many dedicated batch variants plus synthesis wrappers, even though a generic batch_locus_call already exists. While the server integrates many plant-genomics data sources, the surface feels heavy and contains redundancy that could be consolidated. This is too many tools for an agent to navigate comfortably, despite the breadth of the domain.
The surface covers a wide range of plant-genomics needs: locus resolution, sequences, BLAST, GO and plant ontology, pathways, homology, interactions, coexpression, expression, variation, GWAS, structure, domains, literature, and TF motifs. Minor gaps remain, such as no arbitrary genomic-region sequence retrieval, no multiple sequence alignment or phylogenetic tree tool, and several tools restricted to Arabidopsis or a 12-organism list. Still, the read-only research surface is robust overall.
Maintenance
Related MCP Connectors
Protein research over UniProtKB — search by function, fetch curated records, map IDs, proteomes.
Official STRING database MCP server. Query for protein-protein interactions, enrichment, annotations, homology, and PPI networks.
g:Profiler (University of Tartu) — functional enrichment analysis for a gene list against GO…
Link compounds to protein targets, rank bioactivity, and look up drug mechanisms and indications.
Related MCP Servers
- AlicenseAqualityAmaintenanceSearches and fetches research datasets across Zenodo, DataCite (Dryad/Figshare/Dataverse/OSF), NCBI omics archives (GEO/SRA/BioProject), and the literature (PubMed/OpenAIRE) through one normalized model — deduplicating by DOI, expanding organism queries with NCBI Taxonomy synonyms, and bridging papers to the datasets they produced. Resolves citations and open-access full text, and downloads files.66,944 PyPI4MIT
- AlicenseAqualityAmaintenanceGrounds gene-nomenclature work in the HUGO Gene Nomenclature Committee (HGNC) dataset, enabling resolution of gene symbols and IDs to canonical HGNC identifiers, plus cross-references and batch operations.9MIT
- AlicenseNot gradedqualityAmaintenanceRead-only biomedical MCP server connecting PubMed, ClinicalTrials.gov, ClinVar, gnomAD, OncoKB, Reactome, KEGG, UniProt, PharmGKB, CPIC, OpenFDA, Monarch Initiative, GWAS Catalog, and more. One command grammar for all biomedical entities — genes, variants, diseases, drugs, trials, articles, phenotypes, pathways, proteins, diagnostics, and adverse events. 27 tools. Apache-2.0 license.1Apache 2.0
- AlicenseNot gradedqualityCmaintenanceAn MCP server providing public plant bioinformatics APIs including UniProt, NCBI, InterProScan, PDB, AlphaFold, Ensembl Plants, and web-based resources like Sol Genomics and BAR, without local data. It supports gene lookups, protein summaries, structure retrieval, and functional annotations through natural language.MIT