AlphaGenome MCP Server
This server lets Claude agents analyze genetic variants with AlphaGenome, answering single-nucleotide variants instantly from the precomputed Atlas or running live model inference for indels and other variants.
Score one variant's regulatory effects per modality (expression, CAGE, chromatin, histone marks, TF binding, splicing) with scores, calibrated quantiles, and tissue/cell context
Batch-score and rank up to 100 variants, automatically routing each to the Atlas or live inference and reporting the source mix
Scan a genomic region (up to 10,000 bp, or 50,000 bp with opt-in) to rank every substitution without running the model
Look up one or many single-nucleotide variants in the Atlas with default or custom scorers
Report the AlphaGenome Variant Impact (AVI) score plus its 18-feature attributions for SNVs from the Atlas
Assess predicted effect size across modalities (name kept for compatibility; never classifies as pathogenic/benign)
Filter variants by an absolute-quantile threshold (default 0.99) as an effect-size filter
Generate fuller research reports and plain-language explanations of a variant's predicted effects
Predict splice impact, expression impact, TF binding impact, chromatin impact, and allele-specific effects
Compare two variants, protective-vs-risk-labelled variants, alternate alleles at one position, or variants affecting the same gene
Analyze GWAS loci, screen a batch within one modality, compare variants across tissues, or run tissue-specific predictions
List the Atlas scorers (names, track counts, assays) to pass into other tools
Enforce input validation (chromosomes chr1-22/X/Y, 1-based hg38 positions, A/T/G/C alleles, UBERON terms) and cap responses at top_n rows and 40,000 characters
Always state the answer's source (
atlas,live, orlive (atlas fallback)), returning only scores and quantiles94never a clinical classification.
Provides natural language access to Google DeepMind's AlphaGenome variant effect prediction API for analyzing genomic variants, assessing pathogenicity, predicting tissue-specific effects, and evaluating regulatory impacts across 11 molecular modalities including RNA-seq, ChIP-seq, ATAC-seq, and splicing.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AlphaGenome MCP Serveranalyze the pathogenicity of rs429358 in brain tissue"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AlphaGenome MCP Server

AlphaGenome as a tool for Claude agents. A Model Context Protocol (MCP) server that lets an agent turn a researcher's question into an AlphaGenome analysis.
Demo
Ninety seconds in Claude Code, shown at 2x speed, with the published package and the real API. The video has no sound; pause it to read an answer.
https://github.com/user-attachments/assets/cbd083ee-53f9-4b35-b11e-3b33e10f7df9
Time | Prompt | What happens |
0:08 |
| A single-nucleotide variant: answered from the Atlas, with the AVI score |
0:36 |
| The Atlas cannot hold an indel: live inference, and the result says so. A live result has no AVI score |
1:05 |
| 6,003 substitutions ranked from the Atlas without running the model |
The explanations in the video are written by Claude from the tool results. The tools themselves return scores and calibrated quantiles, and never a pathogenicity call.
Related MCP server: EVEE MCP Server
Overview
A researcher asks a question; the agent decides which AlphaGenome calls answer it, runs them, and reads the results back. This server gives the agent the two ways AlphaGenome can be asked:
The AlphaGenome Atlas for single-nucleotide variants. The Atlas holds precomputed predictions for every possible single-nucleotide substitution in the human genome, so a lookup answers in seconds and a whole region can be ranked without running the model.
Live inference for indels and other variants the Atlas cannot precompute.
The server chooses between the two automatically, and every result states its source (source: atlas or source: live), so a precomputed score is never mistaken for a fresh model call or the other way round.
Key Features:
Atlas first: single-nucleotide variants are answered from the precomputed Atlas; live inference runs only when it is needed
Always says where the answer came from:
source: atlas,source: live, orsource: live (atlas fallback: <reason>)One shape for both sources: live inference uses the SDK's
score_variantwith its recommended variant scorers, which have the same names as the Atlas scorers, so an Atlas result and a live result can be read side by sideAVI score and its attributions: the AlphaGenome Variant Impact score (
AVI_SCORE) and its feature attributions are first-class outputs for single-nucleotide variantsRegion scans without the model: rank every substitution in a region by the AVI score
24 tools: 4 Atlas tools and 20 variant tools
These are model predictions for research prioritization, not clinical classifications. Every tool reports AlphaGenome's scores and calibrated quantiles as returned. No tool classifies a variant as pathogenic or benign, assigns a risk label, or states a percent change, on either path. Not for diagnosis or treatment decisions.
⚡ Quick Start
Get started in 3 minutes:
Get an API key at https://deepmind.google.com/science/alphagenome (free for non-commercial use). You need Node.js 18+ and Python 3.10+.
Install dependencies
pip install alphagenome numpyIf this fails, or you are on Windows, use a virtual environment: see Python environment.
Add to your MCP client (supports Claude Code, Claude Desktop, Gemini CLI, Cursor, Windsurf)
claude mcp add alphagenome --env ALPHAGENOME_API_KEY=YOUR_API_KEY -- npx -y @jolab/alphagenome-mcp@latestSee Installation for other MCP clients.
Run your first query
Restart your MCP client and try:
"Use alphagenome to analyze chr19:44908684T>C"View results
This is a single-nucleotide variant, so it is answered from the Atlas in a few seconds and the report starts with
Source: atlas. An indel would run live inference instead (typically 3-10 seconds) and saySource: live.
Want more? Check out the 24 tools below.
Architecture
System Design
AlphaGenome MCP Server implements a multi-tier architecture:
┌─────────────────────────┐
│ Researcher │
└───────────┬─────────────┘
│ Natural language query
↓
┌─────────────────────────┐
│ Claude Desktop │ ← MCP Client
└───────────┬─────────────┘
│ JSON-RPC over stdio
↓
┌─────────────────────────┐
│ MCP Server (TypeScript)│ ← Tool routing, validation
└───────────┬─────────────┘
│ subprocess
↓
┌─────────────────────────┐
│ Python Bridge │ ← Interface to AlphaGenome SDK
└───────────┬─────────────┘
│ HTTP
↓
┌─────────────────────────┐
│ AlphaGenome API │ ← Google DeepMind's service
└─────────────────────────┘Atlas or live inference
AlphaGenome Atlas | Live inference | |
What it answers | Single-nucleotide substitutions on hg38 (chr1-22, chrX, chrY) | Any single variant: single-nucleotide, indel, multi-nucleotide |
How | Looks up precomputed scores | Runs the AlphaGenome model |
Typical time | 2-5 seconds per variant, about 7 seconds for a 2,000 bp region | 3-10 seconds per variant |
Six tools choose between the two on their own: predict_variant_effect, batch_score_variants, assess_pathogenicity, batch_pathogenicity_filter, generate_variant_report and explain_variant_impact. They take an optional source parameter:
| Behaviour |
| Atlas when |
| Atlas only. Never falls back; a variant the Atlas cannot answer is an error |
| Always runs the model |
The fallback is deliberately narrow. It happens only when the Atlas reports that it does not hold the variant. Authentication, rate limit, network and timeout errors are returned as errors, and so is a variant the Atlas rejects: if the reference base does not match hg38, the Atlas says which base it expected, and the server passes that message on instead of running the model on a mistyped variant. In a batch, each variant is routed on its own and the result reports how many came from each source and how many fell back. A mixed batch is reported as two separately ranked groups: the Atlas group is ranked by the AVI score by default, live inference has no AVI score, so one merged ranking would compare different quantities.
Two primitives, one shape
Every tool is a view over two primitives: score one variant and score many variants and rank them. Each is answered by the Atlas or by live inference, and both return the same thing: for each scorer, a matrix of rows (the variant, or one row per gene or junction) by tracks (tissues, cell types, assays), with a calibrated quantile for every score.
Atlas client.query_variant(variant, requested_scorers=[...])
Live inference client.score_variant(interval, variant, RECOMMENDED_VARIANT_SCORERS[...])
-> {scorer: AnnData} -> one shared summarizer -> ranked rows + sourceThe scorer names are the same on both sides (RNA_SEQ, CAGE, DNASE, CHIP_HISTONE, CHIP_TF, SPLICE_SITES, ...). The AVI scorers are the exception: the Atlas alone serves them. For chr17:49210289 C>T the two sources agree on the strongest track of five of the six default scorers, with scores that are close but not identical (DNASE in HeLa-S3: 2.778 from the Atlas, 2.846 live; CAGE in HeLa-S3: 1.866 and 2.121).
A matrix is never returned. One variant is thousands of numbers per scorer, so every result is the strongest cells, ranked, with the gene, tissue and assay of each.
Input Validation
All inputs undergo validation before API submission:
Chromosomes: Pattern-matched for chr1-22, chrX, chrY
Positions: Validated as positive integers
Alleles: A/T/G/C nucleotide validation
Tissue types: UBERON ontology term validation
Invalid inputs return human-readable error messages, enabling conversational error recovery.
Available Tools
AlphaGenome Atlas
Precomputed scores, no model call. Every Atlas response is a summary: ranked rows, never a full score matrix. One variant alone is thousands of numbers per scorer (61 genes x 371 tracks for RNA-seq at the APOE locus), so responses are capped at top_n rows (default 25, maximum 100) and at 40,000 characters.
atlas_list_scorers
The scorers the Atlas serves (22 at the time of writing), with the number of tracks and the assays behind each. Cached for the session.
"Which scorers does the AlphaGenome Atlas have?"atlas_lookup_variant
Scores of one single-nucleotide variant: the strongest tracks per scorer, ranked by absolute score, with the gene, tissue or cell type, and assay of each.
"Look up chr19:44908684 T>C in the AlphaGenome Atlas"atlas_lookup_variants
Up to 500 single-nucleotide variants in one call, ranked by absolute score. Variants the Atlas does not hold, and variants it rejects (for example a reference base that does not match hg38), are listed separately with the reason instead of failing the whole call.
"Rank these 200 GWAS SNPs by their Atlas scores"atlas_scan_region
Every possible single-nucleotide substitution in a region, ranked by absolute score. Answers "which positions in this region matter most?" without running the model. At most 10,000 bp; up to 50,000 bp only with allow_large_region: true.
"Scan chr17:49209289-49211289 and show the 10 substitutions with the largest predicted effect"Default scorers. Override with the scorers parameter; names come from atlas_list_scorers.
Used for | Default | Why |
One variant ( |
| One representative scorer per modality: overall impact, expression, transcription start, accessibility, histone marks, TF binding, splicing |
Many variants and regions ( |
| One number per variant makes the ranking meaningful and keeps the call inside the API's request quota. Measured on a 2,000 bp scan: |
For batch_score_variants, the scoring_metric picks the Atlas scorer when scorers is not given: rna_seq uses RNA_SEQ, splice uses SPLICE_SITES, regulatory_impact and combined use AVI_SCORE.
The AVI score. The AlphaGenome Variant Impact (AVI) score combines AlphaGenome's predictions with AlphaMissense and other features into one number per variant. The Atlas serves it through the API as three scorers, and this server treats them as first-class outputs:
Scorer | What it is |
| The score itself, one value per variant, with its calibrated quantile. The default for ranking many variants and for region scans |
| The contribution of each of the 18 features to the score ( |
| The values of those 18 features |
The AVI score exists for single-nucleotide variants only, because only the Atlas serves it. A live result has no AVI score, and says so rather than approximating one. Scores and quantiles are reported exactly as returned; they are not turned into a pathogenic or benign call.
Limits to know about.
The Atlas API has a requests-per-minute quota, and a region scan uses one request per 32 bp. That is why a scan is limited to 10,000 bp unless
allow_large_regionis set. Long scans wait and retry when the quota is hit. If the time limit (ALPHAGENOME_TIMEOUT_MS) is reached first, the partial result is returned, markedIncomplete, with the range that was really scanned; it is never silently truncated.Scorers with one row per gene or junction (
RNA_SEQ,SPLICE_JUNCTIONS, ...) return more data per request than the API allows for a region scan. Scan withAVI_SCOREfirst, then look up the top variants withatlas_lookup_variant.Human reference genome (hg38) only. Positions are 1-based.
Tools that choose between the Atlas and live inference
All six take source (auto, atlas, live) and scorers.
predict_variant_effect
The strongest tracks of each modality for one variant, with score, quantile, gene, tissue and assay. From the Atlas the AVI score is included. Optional output_types and tissue_type.
"Use alphagenome to analyze chr19:44908684T>C"batch_score_variants
Up to 100 variants, each routed on its own, ranked by predicted effect. Reports how many came from each source and how many fell back.
"Use alphagenome to score these 50 variants and show the top 10"assess_pathogenicity
Predicted effect size across modalities, and the AVI score from the Atlas. The name is kept for compatibility: it does not classify. classification is always null.
"Use alphagenome to assess rs429358"batch_pathogenicity_filter
Keeps the variants whose largest absolute quantile reaches threshold (default 0.99), ranked, per source. A filter on predicted effect size, not on pathogenicity.
"Use alphagenome to keep the variants above the 99.9th percentile"generate_variant_report
A fuller report of one variant: more rows per scorer and, from the Atlas, the AVI score with its feature attributions. A research summary, not a clinical report: no classification, no recommendation.
"Use alphagenome to generate a report for rs429358"explain_variant_impact
Plain sentences that restate the numbers: the AVI score and its largest contributions, then the strongest effect of each modality ordered by quantile, with the direction for signed scorers. Descriptive only.
"Use alphagenome to explain the predicted effect of rs429358"Live-inference tools
These run score_variant and work for single-nucleotide variants, indels and multi-nucleotide variants. Every result states source: live. Tools that take one variant accept an optional tissue_type (a name such as brain, or an ontology CURIE such as UBERON:0000955).
Tool | What it returns |
| Splice sites, splice site usage and splice junctions, with the gene and junction of each |
| RNA-seq (log fold change per gene and tissue) and CAGE |
| TF ChIP-seq, with the factor and cell type of each track |
| ATAC-seq and DNase-seq |
| Alternate against reference allele: |
| Every regulatory modality at once, to see where the predicted effect concentrates |
| The strongest effect of each scorer within each of several tissues |
| Two variants side by side, and which has the larger quantile per scorer |
| The same comparison with the caller's labels; the tool does not judge which allele is protective |
| The alternate alleles of one position, ranked |
| The variants of a locus, ranked |
| Variants ranked by their effect on one gene ( |
| Variants ranked within one modality: |
| One ranking per tissue |
Rankings over several scorers use the largest absolute quantile, because quantiles are calibrated and the raw scores of different scorers are not on one scale. A ranking over one scorer uses its absolute score.
Installation
Requirements
Node.js ≥18.0.0
Python ≥3.10 (required by the
alphagenomepackage)alphagenome≥0.9.0 for the Atlas toolsAlphaGenome API key: https://deepmind.google.com/science/alphagenome (free for non-commercial use)
Python packages:
alphagenome,numpy
Environment variables
Variable | Purpose |
| Your API key. Put it in the |
| Python interpreter to use, for example a virtualenv ( |
| Time allowed for one call, in milliseconds (default 180000). Raise it for large region scans and batches |
Python environment
The server runs the AlphaGenome Python SDK in a subprocess, so it needs a Python 3.10+ interpreter that has alphagenome installed. A virtual environment is the reliable way to get one: it avoids the "externally managed environment" error of system Pythons on macOS and Linux, and on Windows, where python3 usually does not exist and python is often not on the PATH, it gives you a path to point at.
# macOS / Linux
python3 -m venv ~/.alphagenome-venv
~/.alphagenome-venv/bin/pip install alphagenome numpy# Windows (PowerShell)
py -3 -m venv $HOME\.alphagenome-venv
& $HOME\.alphagenome-venv\Scripts\pip install alphagenome numpyThen tell the server which interpreter to use with ALPHAGENOME_PYTHON:
claude mcp add alphagenome \
--env ALPHAGENOME_API_KEY=YOUR_API_KEY \
--env ALPHAGENOME_PYTHON=/Users/you/.alphagenome-venv/bin/python \
-- npx -y @jolab/alphagenome-mcp@latestIn a JSON configuration it goes in the same env block as the key (use the full path; on Windows, double the backslashes):
"env": {
"ALPHAGENOME_API_KEY": "YOUR_API_KEY",
"ALPHAGENOME_PYTHON": "C:\\Users\\you\\.alphagenome-venv\\Scripts\\python.exe"
}If alphagenome is installed for the python3 (or python) on your PATH, you can skip ALPHAGENOME_PYTHON.
Setup
1. Install Python dependencies (or use the virtual environment above):
pip install alphagenome numpy2. Configure for your MCP client:
claude mcp add alphagenome --env ALPHAGENOME_API_KEY=YOUR_API_KEY -- npx -y @jolab/alphagenome-mcp@latestThis writes the server to ~/.claude.json. Add --scope project to write it to .mcp.json in the current project instead.
Test:
"Use alphagenome to analyze chr19:44908684T>C"Add to claude_desktop_config.json:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"alphagenome": {
"command": "npx",
"args": ["-y", "@jolab/alphagenome-mcp@latest"],
"env": { "ALPHAGENOME_API_KEY": "YOUR_API_KEY" }
}
}
}Test:
"Use alphagenome to analyze chr19:44908684T>C"Add to ~/.gemini/settings.json:
{
"mcpServers": {
"alphagenome": {
"command": "npx",
"args": ["-y", "@jolab/alphagenome-mcp@latest"],
"env": { "ALPHAGENOME_API_KEY": "YOUR_API_KEY" }
}
}
}Test:
"Use alphagenome to analyze chr19:44908684T>C"Add to .cursor/mcp.json in your project root:
{
"mcpServers": {
"alphagenome": {
"command": "npx",
"args": ["-y", "@jolab/alphagenome-mcp@latest"],
"env": { "ALPHAGENOME_API_KEY": "YOUR_API_KEY" }
}
}
}Test:
"Use alphagenome to analyze chr19:44908684T>C"Add to your Windsurf settings JSON:
{
"mcpServers": {
"alphagenome": {
"command": "npx",
"args": ["-y", "@jolab/alphagenome-mcp@latest"],
"env": { "ALPHAGENOME_API_KEY": "YOUR_API_KEY" }
}
}
}Test:
"Use alphagenome to analyze chr19:44908684T>C"Verification
Check that the client can start the server. In Claude Code:
claude mcp listalphagenome: ... ✔ Connectedmeans the server starts. It does not yet prove that Python and the key work; the first query does.Restart your MCP client and ask:
"Use alphagenome to analyze chr19:44908684T>C"Expected: a report that starts with
Source: atlaswithin a few seconds. An indel (for examplechr17:49210289 CCC>C) saysSource: liveand typically takes 3-10 seconds.
Troubleshooting
The server reports problems as tool errors and keeps running. The message tells you which case you are in:
Message | Cause and fix |
| No Python on the PATH of the MCP client. Set |
| The interpreter in the message has no |
| Put |
| The key was rejected. Check for a truncated or expired key |
| The |
| The API's per-minute quota. Wait a minute; scan a smaller region |
| A large scan or batch. Raise |
|
|
Tip: say "use alphagenome" in a query when the client does not pick the server on its own.
Usage Examples
All examples are actual results from the API (September 2026) for the two APOE variants that define the ε4 and ε2 alleles.
One variant, from the Atlas
User: "Use alphagenome to assess rs429358" (chr19:44908684 T>C){
"source": "atlas",
"variant": "chr19:44908684:T>C",
"avi_score": { "score": 0.4999, "quantile": 0.9869896 },
"strongest_effects": [
{ "scorer": "RNA_SEQ", "score": -0.038891, "quantile": -0.9996803, "where": "APOE, psoas muscle, polyA plus RNA-seq" },
{ "scorer": "CAGE", "score": 0.077866, "quantile": 0.9774153, "where": "pons, hCAGE" },
{ "scorer": "DNASE", "score": 0.12269, "quantile": 0.9463574, "where": "brain, DNase-seq" },
{ "scorer": "CHIP_HISTONE", "score": 0.13507, "quantile": 0.9953638, "where": "H3K27me3, common myeloid progenitor, CD34-positive, Histone ChIP-seq" },
{ "scorer": "CHIP_TF", "score": -0.16594, "quantile": -0.9972084, "where": "MYC, K562, TF ChIP-seq" },
{ "scorer": "SPLICE_SITES", "score": 0.022699, "quantile": 0.9561927, "where": "APOE" }
],
"largest_abs_quantile": 0.9996803,
"classification": null
}The same variant, explained
User: "Use alphagenome to explain the predicted effect of rs429358"AlphaGenome Variant Impact (AVI) score: 0.4999 (quantile 0.9869896).
Largest contributions to the AVI score: CACTUS_241_WAY (0.2944), PHASTCONS_470_WAY (0.14471), ALPHAMISSENSE (0.029138).
RNA_SEQ: largest predicted effect in APOE, psoas muscle, polyA plus RNA-seq, score -0.038891, quantile -0.9996803 (negative: predicted lower with the alternate allele than the reference).
CHIP_TF: largest predicted effect in MYC, K562, TF ChIP-seq, score -0.16594, quantile -0.9972084 (negative: ...).
...An indel, by live inference
User: "Use alphagenome to analyze the deletion chr17:49210289 CCC>C"**Variant**: chr17:49210289:CCC>C
**Source**: live
**Note**: Live inference because: deletion (CCC>C); the Atlas covers single-nucleotide substitutions only.
## RNA_SEQ
61 row(s) x 371 track(s); largest absolute score 0.8843, median 9.28e-4.
| Score | Quantile | Where |
|---|---|---|
| -0.8843 | -0.9999997 | ABI3, cardiac atrium fibroblast, total RNA-seq |The same SNV from both sources
chr17:49210289 C>T with source: atlas and with source: live, strongest track per scorer:
Scorer | Atlas: score (quantile), where | Live: score (quantile), where |
RNA_SEQ | 0.8612 (0.9999993), ABI3, LHCN-M2 | 0.9677 (0.9999996), ABI3, HFFc6 |
CAGE | 1.866 (0.9999712), HeLa-S3 | 2.121 (0.9999838), HeLa-S3 |
DNASE | 2.778 (0.9998919), HeLa-S3 | 2.846 (0.9999024), HeLa-S3 |
CHIP_HISTONE | 1.846 (0.9999757), H3K27ac, HeLa-S3 | 1.967 (0.9999808), H3K27ac, HeLa-S3 |
CHIP_TF | 1.675 (0.9999024), POLR2A, HeLa-S3 | 1.785 (0.9999371), POLR2A, HeLa-S3 |
SPLICE_SITES | 0.06445 (0.9875146), GNGT2 | 0.04785 (0.9813679), GNGT2 |
The numbers are close, not identical: the Atlas was computed ahead of time and the live call runs the model now.
Two variants side by side
User: "Use alphagenome to compare APOE rs429358 and rs7412"{
"source": "live",
"variants": [
{ "label": "variant1", "variant": "chr19:44908684:T>C",
"effects": [ { "scorer": "RNA_SEQ", "score": -0.037954, "quantile": -0.9996693, "where": "APOE, psoas muscle, polyA plus RNA-seq" }, "..." ] },
{ "label": "variant2", "variant": "chr19:44908822:C>T",
"effects": [ { "scorer": "RNA_SEQ", "score": 0.035793, "quantile": 0.9997013, "where": "APOE, endothelial cell of umbilical vein, polyA plus RNA-seq" }, "..." ] }
],
"larger_abs_quantile_by_scorer": { "RNA_SEQ": "variant2", "...": "..." }
}Tissue-specific
User: "Use alphagenome to compare rs429358 in brain and liver"In brain the strongest RNA-seq effect is on APOE (score -0.0084, quantile -0.9943216); DNase-seq in brain is 0.1224 (quantile 0.9463574) and in liver -0.004807 (quantile -0.1842174).
Filtering a mixed batch
User: "Use alphagenome to keep the variants above the 99th percentile: rs429358, rs7412, and the deletion chr17:49210289 CCC>C"The two SNVs are answered from the Atlas and ranked by the AVI score (rs7412: 0.8636, quantile 0.995461, kept; rs429358: quantile 0.9869896, not kept). The deletion is answered by live inference (largest quantile 0.9999997, kept). The result reports the two groups separately: "source": "mixed", "answered_from": { "atlas": 2, "live": 1, "atlas_fallback": 0 }.
Performance
Atlas: 2-5 seconds per variant, about 7 seconds for a 2,000 bp region scan (with
AVI_SCORE)Live inference: typically 3-10 seconds per variant (measured through the server, September 2026; depends on API load)
Modalities: 11 (RNA-seq, CAGE, PRO-cap, splice sites, DNase, ATAC, histone mods, TF binding, contact maps)
Development
Build from Source
git clone https://github.com/taehojo/alphagenome-mcp.git
cd alphagenome-mcp
npm install
pip install -r requirements.txt
npm run buildProject Structure
src/
├── index.ts # MCP server entry point
├── alphagenome-client.ts # API client (Python bridge, timeout, interpreter choice)
├── routing.ts # Atlas or live inference: the rule, as pure functions
├── variant-tools.ts # The 20 variant tools, over two primitives and a narrow backend interface
├── tools.ts # MCP tool definitions
├── types.ts # TypeScript type definitions
├── tests/ # Unit tests (no API key needed)
└── utils/
├── config.ts # Environment variables
├── validation.ts # Input validation (Zod schemas)
└── formatting.ts # One set of formatters for both sources, and the size cap
scripts/
├── alphagenome_bridge.py # Bridge: one JSON request in, one JSON response out
├── atlas_actions.py # Atlas lookups and region scans
├── live_actions.py # Live inference with score_variant
├── summaries.py # The summarizer shared by both sources (numpy only)
└── tests/ # Python unit tests for the summarizerTesting
npm test # Build, then run the TypeScript unit tests (no API key needed)
npm run test:python # Python unit tests for the shared summarizer (numpy and pandas only)
npm run docs:api # Regenerate docs/API.md from the tool definitions in src/tools.ts
npm run lint # ESLint check
npm run typecheck # TypeScript type checking
npm run build # Compile to build/Roadmap
Not in this release, and not promised by any tool above:
Combinations of variants: scoring several variants together on one haplotype, rather than one at a time.
Custom sequences: predictions for a sequence the caller supplies, rather than a variant on the reference genome.
Both need live inference and neither can be precomputed, so they fit the same design: the Atlas where it can answer, the model where it cannot, and the source on every result.
Why this server exists
From a microglia gene to a variant map in about 60 seconds. The research workflow this server came from: which non-coding variants near ABI3 change its expression, answered with a traditional pipeline and with one sentence in Claude Code. Click the image for full resolution.

AlphaGenome at study scale
Rare variants and AlphaGenome-predicted regulatory impact in 85 Alzheimer's disease genes: an interactive browser of the rare variants (MAF < 1%) in 85 AD-associated genes from ADSP whole-genome sequencing, scored with AlphaGenome across eight regulatory modalities and compared with case-control allele frequencies in 24,595 ADSP R4 participants, with an independent assessment in 11,545 ADSP R5 participants (Jo et al., manuscript under review). Code and the 9,943-variant analysis table: taehojo/rarevariants.
The scores in that study were computed with the AlphaGenome Python SDK directly, not with this server, and before its 0.3.0 scoring. They are not output of this server.
Citation
If you use this software in your research, please cite:
@software{jo2025alphagenome_mcp,
author = {Jo, Taeho},
title = {AlphaGenome MCP Server},
year = {2025},
url = {https://github.com/taehojo/alphagenome-mcp},
version = {0.2.0}
}AlphaGenome model:
@article{avsec2026alphagenome,
title = {Advancing regulatory variant effect prediction with {AlphaGenome}},
author = {Avsec, Žiga and Latysheva, Natasha and Cheng, Jun and others},
journal = {Nature},
volume = {649},
number = {8099},
pages = {1206--1218},
year = {2026},
doi = {10.1038/s41586-025-10014-0}
}Acknowledgments
Google DeepMind for developing and providing access to the AlphaGenome API
Anthropic for developing the Model Context Protocol specification and Claude Desktop
License
MIT License - Copyright (c) 2025 Taeho Jo
See LICENSE file for details.
Links
npm Package: https://www.npmjs.com/package/@jolab/alphagenome-mcp
GitHub Repository: https://github.com/taehojo/alphagenome-mcp
AlphaGenome: https://deepmind.google/discover/blog/alphagenome/
Model Context Protocol: https://modelcontextprotocol.io/
Claude Desktop: https://claude.ai/download
Available Tools
24 toolsanalyze_gwas_locusA
Rank the variants of a locus by predicted effect (largest absolute quantile across modalities), to prioritize candidates for follow-up. For single-nucleotide variants only, atlas_lookup_variants or atlas_scan_region is faster and adds the AVI score.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | Optional: locus end, for the label | |
| start | No | Optional: locus start, for the label | |
| variants | Yes | Variants to score (1-100) | |
| chromosome | No | Optional: locus chromosome, for the label |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that inference runs live via score_variant with recommended scorers, that the result is tagged source: live, that it covers SNVs/indels/MNVs, and that outputs are research predictions with no pathogenic/benign call. It does not mention cost, rate limits, or latency for a 1-100 variant live call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, purpose and alternative front-loaded before the mechanism and disclaimer. Mostly efficient, though the SNV phrasing in the first paragraph and the modality coverage in the second overlap enough to introduce mild ambiguity rather than pure redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, so the description must supply behavioral context, and it covers the live-inference source flag, modality scope, and non-clinical disclaimer. It stops short of describing the shape of the ranking output or how many results are returned, which would fully complete it for a 1-100 variant tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so start/end/chromosome/variants are already well documented including the 1-100 bound. The description adds no additional semantics (e.g., interpretation of the locus label fields) beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (rank) and resource (variants of a locus) plus the ranking criterion (largest absolute quantile across modalities) and the goal (prioritize follow-up candidates). An agent can distinguish it from sibling scorers like predict_variant_effect or batch_score_variants from this alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes SNV-only inputs to atlas_lookup_variants or atlas_scan_region, naming the condition (SNV-only) and the benefit (faster, adds AVI). However the phrasing 'For single-nucleotide variants only' sits awkwardly beside the later claim that the tool handles SNVs, indels and MNVs, leaving the exact when-not boundary slightly muddy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
annotate_regulatory_contextA
The predicted effects of a variant across every regulatory modality at once: accessibility, histone marks, TF binding, CAGE, RNA-seq, splice sites, polyadenylation and contact maps. Shows where the predicted effect concentrates; it does not label the variant.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that live inference runs via score_variant using SDK-recommended scorers, the variant types supported (SNV, indel, MNV), that the response carries `source: live`, and that it makes no pathogenic/benign call. This is rich behavioral context an agent cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded paragraphs: what it predicts, how it runs, and the disclaimers. Efficient overall, though the final disclaimer paragraph is somewhat long and repeats the non-clinical framing already implied by the first paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains what the result contains (source marker, scores and calibrated quantiles) and what it deliberately omits (pathogenicity calls). It doesn't detail the per-modality result structure, but for an annotation tool with full schema coverage the coverage is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter (chromosome, position, ref, alt, tissue_type) is already documented in the schema with patterns and defaults. The description adds no syntax or semantics beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (annotate/predict) and resource (variant effects across all regulatory modalities), enumerating the modalities covered: accessibility, histone marks, TF binding, CAGE, RNA-seq, splice sites, polyadenylation and contact maps. The phrase 'across every regulatory modality at once' implicitly differentiates it from single-modality siblings like predict_splice_impact or predict_tf_binding_impact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the intended use ('research prioritization, not clinical classifications') and states it does not label the variant, which steers an agent away from it when a pathogenic/benign call is needed (assess_pathogenicity). It does not, however, name a specific sibling alternative or give explicit when-not conditions beyond the clinical boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assess_pathogenicityA
Predicted effect size of a variant across modalities, for prioritization. The tool name is kept for compatibility: it does NOT classify a variant as pathogenic or benign, and classification is always null.
Returns the strongest effect per scorer (score, calibrated quantile, where it was seen), the largest absolute quantile, and, for a single-nucleotide variant answered from the Atlas, the AlphaGenome Variant Impact (AVI) score. avi_score is null on the live path, because the AVI score is served by the Atlas only.
Source: a single-nucleotide variant is answered from the precomputed AlphaGenome Atlas; an indel or multi-nucleotide variant runs live inference (score_variant). Both return the same scorers in the same shape. Chosen automatically, overridable with source, and always stated in the result.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "How large is the predicted effect of chr19:44908684 T>C?"
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| source | No | Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered. | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that `classification` is always null, that `avi_score` is null on the live path because only the Atlas serves it, exactly which source answers which variant class, that the chosen source is always echoed in the result, and the shape of the returned per-scorer data. It also sets a clear research-use boundary (no pathogenic/benign call).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the disambiguation and return shape, and each paragraph has a distinct job (naming, returns, source routing, caveat). It is slightly long and repeats the no-pathogenic-call point twice (once after the name caveat, again in the closing paragraph), which is minor redundancy rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no annotations and no output schema, the description is largely sufficient: it explains the return contents (strongest effect per scorer, largest absolute quantile, AVI) and the source semantics. It stops short of a fully enumerated result structure and does not note any cost or latency difference for the live-inference path, which an agent weighing `source` would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters including `source`, `scorers`, and `tissue_type`. The description adds only marginal parameter meaning beyond that (the AVI-scorers-only-in-Atlas caveat, which is already in the schema), so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource ('predicted effect size of a variant across modalities') and the second immediately neutralizes the misleading tool name by stating it does NOT classify pathogenicity and that `classification` is always null. It also delineates the two answering paths (SNV via Atlas, indel/MNV via live inference), so an agent can tell what the tool will and won't produce without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the purpose context ('for prioritization') and explains the automatic source selection plus the `source` override, including what happens on fallback. It does not, however, name or exclude the very similar sibling `predict_variant_effect` or `predict_tissue_specific`, so the agent must infer routing between closely related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
atlas_list_scorersA
List the variant scorers available in the AlphaGenome Atlas, including the AlphaGenome Variant Impact score (AVI_SCORE) and its feature attributions (AVI_SCORE_FEATURE_IMPORTANCE, AVI_SCORE_MODEL_FEATURES).
Returns each scorer's name, number of tracks and the assays behind it. Use these names in the scorers parameter of the other tools. Every scorer except the AVI ones is also available from live inference under the same name. Cached for the session after the first call.
The Atlas holds precomputed AlphaGenome scores for single-nucleotide substitutions on the human reference genome (hg38, chr1-22, chrX, chrY). Indels and multi-nucleotide variants are not in it; use predict_variant_effect for those.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full load and does well: it discloses the return contents (name, number of tracks, assays), a caching behavior ('Cached for the session after the first call'), and the data scope (hg38, chr1-22, chrX, chrY, SNVs only). It does not state permissions or pagination, but for a static enumerator that is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and return contents, then scope constraints. Three short paragraphs earn their place, though the long parenthetical scorer-name enumeration adds bulk that could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param enumerator with no output schema, the description covers what is returned, how to use it downstream, caching, and the dataset boundaries. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description adds cross-tool value by explaining how the returned names feed the `scorers` parameter elsewhere, which is meaningful context beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the variant scorers available in the AlphaGenome Atlas') and enumerates the concrete scorer names returned. An agent can distinguish this enumeration tool from scoring tools like predict_variant_effect without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: 'Use these names in the `scorers` parameter of the other tools,' and names the alternative and exclusion ('Indels and multi-nucleotide variants are not in it; use predict_variant_effect for those'). Both when-to-use and when-not-to-use are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
atlas_lookup_variantA
Look up the precomputed AlphaGenome scores of one single-nucleotide variant. No model call, so it answers in seconds.
Returns, per scorer, the strongest tracks for the variant ranked by absolute score, each with its calibrated quantile, gene, tissue or cell type, and assay. Default scorers: AVI_SCORE plus one per modality (RNA_SEQ, CAGE, DNASE, CHIP_HISTONE, CHIP_TF, SPLICE_SITES).
If the reference base does not match hg38, the Atlas says which base it expected and that message is returned as a validation error.
The Atlas holds precomputed AlphaGenome scores for single-nucleotide substitutions on the human reference genome (hg38, chr1-22, chrX, chrY). Indels and multi-nucleotide variants are not in it; use predict_variant_effect for those.
The response is a summary, never a full score matrix: ranked rows only, capped at top_n (default 25, max 100) and at 40,000 characters.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Look up chr19:44908684 T>C in the AlphaGenome Atlas"
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate base (one of A, C, G, T) | |
| ref | Yes | Reference base (one of A, C, G, T). Must match hg38 at this position. | |
| top_n | No | Rows to return (default: 25, max: 100) | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it does a lot: read-only lookup semantics, no model call, second-scale latency, top_n default/max, a 40,000-character response cap, and explicit validation-error behavior when ref does not match hg38. It omits auth/permission requirements and any rate-limit note, which are the only remaining gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what and how fast, then response shape, then scope limits, then caveats, then an example. Structure is good and most sentences earn their place, though the trailing example sentence adds little given the clear parameter docs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain returns, and it does: per-scorer ranked rows with quantile, gene, tissue/cell type, and assay, plus the summary-not-matrix cap. Combined with scope limits and error behavior, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: it names the default scorer set (AVI_SCORE plus one per modality) and lists the actual modalities, and reaffirms the top_n default/max. It goes beyond restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up), resource (precomputed AlphaGenome scores), and scope (one single-nucleotide variant). It also distinguishes itself from the nearest functional sibling by explaining the Atlas is precomputed-only, so the agent can tell it apart without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says indels and multi-nucleotide variants are not supported and routes those to predict_variant_effect, and implies fast use via 'no model call, answers in seconds'. It does not, however, address when to use this versus the near-identically named atlas_lookup_variants, which is a real ambiguity left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
atlas_lookup_variantsA
Look up precomputed AlphaGenome scores for up to 500 single-nucleotide variants in one call and rank them.
Returns one row per variant (its strongest score and where it was seen), ranked. Default scorer: AVI_SCORE (AlphaGenome Variant Impact), one number per variant. With several scorers the ranking uses the largest absolute quantile. Variants the Atlas does not hold and variants it rejects (for example a reference base that does not match hg38) are listed separately with the reason; they do not fail the call.
The Atlas holds precomputed AlphaGenome scores for single-nucleotide substitutions on the human reference genome (hg38, chr1-22, chrX, chrY). Indels and multi-nucleotide variants are not in it; use predict_variant_effect for those.
The response is a summary, never a full score matrix: ranked rows only, capped at top_n (default 25, max 100) and at 40,000 characters.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Rank these 200 GWAS SNPs by their Atlas scores"
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | Rows to return (default: 25, max: 100) | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| variants | Yes | Single-nucleotide variants to look up (1-500) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: one row per variant, default AVI_SCORE, largest-absolute-quantile ranking with multiple scorers, missing/rejected variants listed separately with reasons rather than failing the call, the 40,000-character response cap, and a research-only (non-clinical) disclaimer. This is well beyond what the schema conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and result shape, then scope limits, then caveats; each paragraph is purposeful. Slightly long, with the closing example and dual disclaimers bordering on extras, but nothing misleading or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain returns, and it does: ranked rows only, capped at top_n and 40,000 characters, with out-of-Atlas/rejected variants reported separately. An agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: top_n defaults to 25 and caps at 100, and the choice of scorers changes the ranking method (largest absolute quantile). The only gap is that scoring syntax/format details remain in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up precomputed scores), resource (AlphaGenome SNV scores), and scope (up to 500 in one call, ranked). It distinguishes itself from siblings by noting indels/MNVs are out of scope and belong to predict_variant_effect, and the singular sibling atlas_lookup_variant is implicitly contrasted by the batch framing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative explicitly for the out-of-scope case ('Indels and multi-nucleotide variants are not in it; use predict_variant_effect for those') and points to atlas_list_scorers as the source of valid scorer names. Both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
atlas_scan_regionA
Scan a genomic region in the AlphaGenome Atlas: every possible single-nucleotide substitution in the interval, ranked. Answers "which positions in this region matter most?" without running the model.
Region width: at most 10,000 bp. Up to 50,000 bp only with allow_large_region=true. A scan is one API request per 32 bp under a requests-per-minute quota, so a large scan can take minutes. If the quota or the time limit stops a scan early, the partial result is returned, marked "Incomplete", with the range that was really scanned; the ranking then covers that range only.
Default scorer: AVI_SCORE (about 7 seconds for 2,000 bp). Multi-track scorers are much slower (2,000 bp: DNASE 17 s, CHIP_TF 86 s). Scorers with one row per gene or junction (RNA_SEQ, SPLICE_JUNCTIONS, ...) cannot be used for a scan: scan with AVI_SCORE, then use atlas_lookup_variant on the top variants.
The Atlas holds precomputed AlphaGenome scores for single-nucleotide substitutions on the human reference genome (hg38, chr1-22, chrX, chrY). Indels and multi-nucleotide variants are not in it; use predict_variant_effect for those.
The response is a summary, never a full score matrix: ranked rows only, capped at top_n (default 25, max 100) and at 40,000 characters.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Scan chr17:49209289-49211289 and show the 10 substitutions with the largest predicted effect"
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | End position (greater than start; at most start + 10,000, or start + 50,000 with allow_large_region) | |
| start | Yes | Start position (1-based, hg38) | |
| top_n | No | Rows to return (default: 25, max: 100) | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| allow_large_region | No | Set to true to scan more than 10,000 bp (up to 50,000 bp). Slower, and the result may be incomplete (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: 10,000 bp cap (50,000 with allow_large_region), one API request per 32 bp under a requests-per-minute quota, minutes-long runtime, partial results marked "Incomplete" with the actually-scanned range, and scorer-dependent timings (AVI ~7 s vs DNASE 17 s vs CHIP_TF 86 s for 2,000 bp). It also caps and characterizes the response and disclaims clinical use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then layers limits, timing, scorer rules, and response shape, with an example at the end. Dense and mostly waste-free, though at roughly 280 words it is longer than strictly necessary and could merge the timing and quota sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return contract itself: ranked rows only (never a full matrix), capped at top_n (default 25, max 100) and 40,000 characters, with "Incomplete" partials and their real range. Together with genome-build scope (hg38) and the research-only caveat, nothing needed to invoke or interpret the call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds real meaning beyond it, notably that scorers must come from atlas_list_scorers, that AVI scorers are Atlas-only, and that allow_large_region trades speed for possible incompleteness. It does not add much on start/end/top_n beyond what the schema already spells out.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb (scan) and resource (genomic region / all SNVs in an interval) and frames the question it answers ("which positions in this region matter most?" without running the model). It explicitly distinguishes itself from siblings atlas_lookup_variant and predict_variant_effect, so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use and when-not: indels/MNV must go to predict_variant_effect, and multi-row scorers like RNA_SEQ cannot be scanned — instead scan with AVI_SCORE and follow up with atlas_lookup_variant on top variants. Alternatives and the conditions selecting them are named directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_modality_screenA
Rank several variants by predicted effect within one modality: expression (RNA_SEQ, CAGE), splicing (SPLICE_SITES, SPLICE_SITE_USAGE), tf_binding (CHIP_TF) or chromatin (DNASE, ATAC).
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| modality | Yes | ||
| variants | Yes | Variants to score (1-100) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden well: it discloses live inference execution, supported variant classes (SNV, indel, MNV), the output marker (source: live), and an explicit research-only / non-clinical limitation. It omits operational traits like latency, cost, or whether anything is persisted, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, front-loaded with the core action and modality/scorer mapping, then execution detail, then the caveat. Every sentence contributes; only mild redundancy in the closing 'as returned' phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, the description covers what it computes and flags output provenance, but leaves the selection problem unsolved — an agent facing ~20 sibling scoring tools has no basis to choose this one. The core semantics are adequate; the comparative context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, and the description compensates by mapping each modality enum value to concrete scorers (RNA_SEQ/CAGE, SPLICE_SITES, CHIP_TF, DNASE/ATAC) — meaning the schema does not carry. It also frames which variant types the variants array accepts, though it adds nothing about the positional/allele constraints already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (rank) applied to a resource (variants) with an explicit scope (within one modality), and the modality-level framing plus scorer mapping is distinctive. However, it never names or distinguishes itself from near-identical siblings such as batch_score_variants, predict_variant_effect, or compare_variants, so the agent cannot tell from the description alone which to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative routing is given despite a very crowded sibling set of scoring/prediction tools. The only routing-like sentence ('score_variant with the SDK's recommended variant scorers') describes internal implementation rather than guiding tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_pathogenicity_filterA
Keep the variants whose predicted effect reaches a threshold, ranked. The tool name is kept for compatibility: it filters on predicted effect size, NOT on pathogenicity, and classifies nothing.
threshold is the smallest absolute calibrated quantile (0 to 1) a variant must reach to be kept (default 0.99): the AVI score's quantile for variants answered from the Atlas, the largest quantile across modalities for live inference. Variants are routed per variant and reported per source; groups from different sources are not comparable.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Which of these variants have a predicted effect above the 99.9th percentile?"
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered. | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| variants | Yes | Variants to score (1-100) | |
| threshold | No | Smallest absolute quantile to keep, between 0 and 1 (default: 0.99) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it explains that variants route per variant, results report per source, and groups from different sources are not comparable. It also discloses that no pathogenic/benign call is made and that scores/quantiles are returned as-is, though it does not cover error or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the key name-correction caveat, then threshold semantics, then the disclaimer and an example. Mostly efficient, though the 'not clinical classifications' disclaimer is restated twice ('no pathogenic/benign call is made'), which is slight redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does supply return context: results are ranked, reported per source, and scores and calibrated quantiles are returned as-is. Combined with the schema-documented source routing, an agent has enough to call and interpret the tool, though pagination/volume of returns is not addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3, but the description adds genuine meaning to `threshold` beyond the schema: the smallest absolute calibrated quantile for Atlas-answered variants versus the largest quantile across modalities for live inference. That source-dependent semantics is not available from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: it keeps variants whose predicted effect reaches a threshold, ranked. It explicitly corrects its misleading name ('filters on predicted effect size, NOT on pathogenicity, and classifies nothing'), which is exactly the disambiguation an agent needs against siblings like assess_pathogenicity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It makes the use context clear (effect-size prioritization, not clinical classification) and gives a concrete triggering question as an example. It stops short of naming an alternative tool for the pathogenicity-classification use case it rules out, leaving that route implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_score_variantsA
Score up to 100 variants and rank them by predicted effect.
Each variant is routed on its own: single-nucleotide variants to the Atlas, the rest to live inference. A mixed batch comes back as two separately ranked groups, because the Atlas group is ranked by the AVI score and live inference has no AVI score; the two must not be compared. The result reports how many variants came from each source and how many fell back.
Scoring metric (used when scorers is not given): rna_seq = RNA_SEQ, splice = SPLICE_SITES, regulatory_impact and combined = AVI_SCORE from the Atlas and every modality (ranked by the largest absolute quantile) from live inference.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Score these 50 variants and show me the top 10 by predicted effect"
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | Variants to return per group (default: 10, max: 100) | |
| source | No | Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered. | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| variants | Yes | Variants to score (1-100) | |
| scoring_metric | Yes | What to rank by |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses key non-obvious behaviors: per-variant routing, two non-comparable ranked groups (AVI vs live), fallback counts in the result, and that outputs are predictions not clinical classifications. It does not cover rate limits, permissions, or error behavior for atlas-only mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose followed by routing rules, metric mapping, and a disclaimer. Four paragraphs are justified by the genuinely complex behavior, though the example at the end is somewhat redundant given the clear opening sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param tool with no output schema and no annotations, the description covers routing, grouping semantics, metric defaults, and the research-use limitation. Missing pieces are minor: no explanation of output fields beyond source counts, and no explicit prerequisite guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents each parameter, and the description's metric mapping (rna_seq → RNA_SEQ, etc.) largely restates the enum. The description adds marginal value by clarifying that scorers come from atlas_list_scorers and that AVI scorers are Atlas-only, but this is mostly schema-adjacent rather than new semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource+scope: 'Score up to 100 variants and rank them by predicted effect.' The batch nature and the 100-variant cap distinguish it from single-variant siblings, and the routing detail (Atlas vs live inference) further differentiates it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains internal routing behavior and the two-group result, but does not state when to use this over sibling batch tools like batch_pathogenicity_filter or batch_modality_screen, nor when a user should prefer single-variant predict_variant_effect. Usage is implied by the batch scope rather than explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_tissue_comparisonA
Rank several variants by predicted effect within each of several tissues: one ranking per tissue.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| tissues | Yes | Tissue names or ontology CURIEs (1-10) | |
| variants | Yes | Variants to score (1-100) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does substantial work: it discloses that inference is live, that the result carries 'source: live', and that outputs are model predictions for research prioritization rather than pathogenic/benign calls. It omits auth requirements, cost/latency expectations for live scoring, and rate limits, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core operation in the first sentence, followed by scope, execution mode, and a necessary caveat. Every sentence earns its place, though the caveat paragraph is slightly verbose for the amount of information conveyed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey the return shape, and it does: one ranking per tissue, live source marker, and calibrated quantiles as returned. The only meaningful gap is the absence of any note on response size or runtime for up to 100 variants across 10 tissues in a live-inference call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only two top-level parameters, so the schema already documents tissues, variants, and every nested allele field. The description adds the implicit pairing semantics (each variant ranked in each tissue) but no formats or constraints beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and output shape: 'Rank several variants by predicted effect within each of several tissues: one ranking per tissue.' The variants-by-tissues matrix framing clearly distinguishes it from flat siblings like batch_score_variants or compare_variants. It does not name a sibling explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives input-eligibility guidance ('works for single-nucleotide variants, indels and multi-nucleotide variants'), which is useful, but never says when to choose this over batch_score_variants, predict_tissue_specific, or compare_variants. Usage is implied by the matrix output rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_allelesA
Rank the alternate alleles of one position by predicted effect.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| alts | Yes | Alternate alleles (1-20) | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that inference is live (and that results carry `source: live`), that outputs are model predictions for research prioritization, and that no pathogenic/benign call is made. It stops short of stating latency, cost, or whether repeated calls are deterministic, which would be valuable for a live-inference tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in the first sentence, then separates mechanics from the clinical-use caveat. Every sentence carries information; the only minor cost is the parenthetical about score_variant, which reads more like implementation detail than agent-facing guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description covers the essentials an agent needs: what is ranked, that inference is live, the supported variant classes, and the interpretation limits. Missing only operational details (runtime, determinism, error behavior) that would complete the picture for a live-inference call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so chromosome, position, ref, and alts are already documented, including patterns and the 1-20 alt limit. The description adds only implicit semantics (that indels/MNVs are expressed through ref/alts), so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Rank the alternate alleles of one position by predicted effect'), which makes the operation clear. It does not explicitly differentiate itself from close siblings like compare_variants or compare_variants_same_gene, so an agent must infer the boundary from the names alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives useful scope guidance by naming the variant classes it handles ('single-nucleotide variants, indels and multi-nucleotide variants') and clarifies that inference is live. However, it never says when to prefer this tool over predict_variant_effect, compare_variants, or batch_score_variants, which is the main routing question for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_protective_riskA
Two variants side by side, labelled as the caller names them. The tool compares predicted effect sizes per modality; it does not judge which allele is protective or a risk.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| risk_variant | Yes | ||
| protective_variant | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that it runs live inference via score_variant with the SDK's recommended scorers, that the result carries `source: live`, that it accepts SNVs, indels and MNVs, and that outputs are AlphaGenome model predictions with no pathogenic/benign call. It does not cover cost, latency, or rate limits of live inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, front-loaded with what the tool does and its key caveat, with no filler. The live-inference note and the research-use disclaimer each occupy one tight sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-nested-parameter tool with no annotations and no output schema, the description covers behaviour (live inference, source marker), accepted variant types, and the interpretive limits of the output. An agent has enough to call it correctly, though a hint about which sibling to prefer would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The visible schema documents each nested field (chromosome, position, ref, alt, variant_id), and the description adds semantics the schema cannot: the protective_variant/risk_variant labels are supplied by the caller and are not validated or judged by the tool. That meaningfully guides how the two required objects should be assigned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it compares two caller-labelled variants and reports predicted effect sizes per modality, and explicitly clarifies it does not assign protective/risk direction. This is clear, though it never names or contrasts with close siblings such as compare_variants, compare_alleles, or compare_variants_same_gene.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one implicit when-not: it does not judge which allele is protective or a risk, which tells the agent not to use it for classification. However, with 20+ siblings including several near-identical comparison tools, no alternative is named and no condition for selecting this over them is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_variantsA
Two variants side by side: the strongest predicted effect of each modality for both, and which of the two has the larger absolute quantile per scorer. A comparison of predicted effect sizes, not of severity.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Compare APOE rs429358 and rs7412"
| Name | Required | Description | Default |
|---|---|---|---|
| variant1 | Yes | ||
| variant2 | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely does: it discloses that this performs live inference via score_variant with the SDK's recommended scorers, that results carry `source: live`, and that outputs are model predictions for research prioritization with no pathogenic/benign call. That is meaningful behavioral context beyond a bare 'compare' statement, though it doesn't quantify latency/cost of live inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core comparison semantics are front-loaded in the first sentence, followed by execution model, applicability, provenance, and a caveat. The 'comparison of predicted effect sizes, not of severity' idea is restated near the end, which is mildly redundant but not wasteful enough to hurt much.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-nested-object tool with no output schema and no annotations, the description covers execution model, accepted variant types, provenance flag, and interpretive limits. It could say more about what the returned per-scorer structure looks like, but the essentials an agent needs to invoke it correctly are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no field-level meaning for variant1/variant2 beyond noting the accepted variant classes (SNV, indel, MNV) and an example using rsIDs. The nested schema properties do carry descriptions, so variant structure is largely covered there, making 3 the appropriate baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete operation on a specific resource: pairwise comparison of two variants, reporting per-modality predicted effects and which has the larger absolute quantile per scorer. It even disambiguates the semantics ('comparison of predicted effect sizes, not of severity'). It stops short of explicitly routing against near-name siblings like compare_variants_same_gene or compare_alleles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the operating context (runs live inference, works for SNVs/indels/MNVs) and gives a worked example, which implies when it applies. However, there is no explicit when-not guidance or pointer to alternatives such as batch_score_variants for many variants or compare_variants_same_gene for the same-gene case, leaving sibling selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_variants_same_geneA
Rank several variants by their predicted effect on one gene. With gene_name, the gene-level scorers (RNA_SEQ, SPLICE_SITES) are restricted to that gene.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| variants | Yes | Variants to score (1-100) | |
| gene_name | No | Optional: gene symbol (e.g., APOE) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses that inference is live, which variant types are supported (SNV, indel, MNV), that output carries source: live, and that results are research-grade model predictions with no clinical call. It omits operational traits like latency, rate limits, or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action before the gene-name behavior and the output/disclaimer detail. Every sentence contributes; the structure is clean and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description usefully covers return semantics (ranking by predicted effect, source marker, calibrated quantiles reported as returned) and covers behavior in the absence of annotations. It could say more about the exact shape/ordering of the ranking result, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema by explaining that gene_name restricts the gene-level scorers (RNA_SEQ, SPLICE_SITES) to that gene, which the schema field ('Optional: gene symbol') does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Rank several variants by their predicted effect on one gene') and adds technical scope (gene-level scorers RNA_SEQ, SPLICE_SITES). It does not explicitly name or distinguish itself from close siblings like compare_variants or compare_alleles, so an agent still has to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the tool is scoped to ranking several variants within a single gene, and the gene_name clause hints at when the gene restriction applies. There is no explicit when-to-use, when-not-to-use, or routing to any of the many comparator siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_variant_impactA
Plain sentences that restate a variant's predicted effects: the AVI score and its largest contributions (from the Atlas), then the strongest effect of each modality ordered by absolute quantile, with the direction for signed scorers. Descriptive only: the sentences restate returned numbers and make no statement about pathogenicity.
Source: a single-nucleotide variant is answered from the precomputed AlphaGenome Atlas; an indel or multi-nucleotide variant runs live inference (score_variant). Both return the same scorers in the same shape. Chosen automatically, overridable with source, and always stated in the result.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Explain the predicted effect of chr17:49210289 C>T in plain language"
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| source | No | Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered. | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and largely succeeds: it explains the auto/atlas/live dispatch, that atlas errors instead of falling back, that live always runs inference, that the source is always reported, and that no pathogenicity call is made and results are research-only. It does not cover latency, rate limits, or failure modes of live inference, but the safety/dispatch profile is well documented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is front-loaded (output shape first, then source routing, then disclaimers, then an example) and each paragraph earns its place. It is somewhat verbose and repeats the source-selection semantics that the schema already states, which is the only real redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly carries the return-value burden by describing what the sentences contain and that the answering source is always stated. Combined with the parameter docs, an agent has enough to call and interpret the tool; only inter-tool routing against its many siblings remains unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all seven parameters are already documented, including the source enum and the note that scorer names come from atlas_list_scorers. The description restates the source behavior but adds no format, syntax, or defaulting detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a concrete verb and resource: it produces plain sentences restating a variant's predicted effects, and spells out exactly what those sentences contain (AVI score, biggest contributions, per-modality strongest effect with direction). It also carves out a distinct niche ('descriptive only... no statement about pathogenicity'), which separates it from assess_pathogenicity and related siblings. However, it never distinguishes itself by name from the many predict_* siblings that also report variant effects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated. The description explains how the source is chosen and that it can be overridden, and gives an example prompt, but it never says when an agent should pick this narration tool over predict_variant_effect, generate_variant_report, or assess_pathogenicity. No explicit when-not guidance is offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_variant_reportA
A fuller report of one variant's predicted molecular effects: more rows per scorer than predict_variant_effect and, from the Atlas, the AVI score with its feature attributions (AVI_SCORE_FEATURE_IMPORTANCE). It is a research summary, not a clinical report: it contains no pathogenicity classification and no recommendation.
Source: a single-nucleotide variant is answered from the precomputed AlphaGenome Atlas; an indel or multi-nucleotide variant runs live inference (score_variant). Both return the same scorers in the same shape. Chosen automatically, overridable with source, and always stated in the result.
The response is a summary, never a full score matrix: ranked rows only, capped at top_n (default 25, max 100) and at 40,000 characters.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Generate a report for chr19:44908684 T>C"
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| source | No | Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered. | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that the response is a capped summary (top_n default 25, max 100; 40,000-character ceiling), that the answering source is always reported, that atlas-only mode errors rather than falls back, and that no pathogenic/benign call is made. These are exactly the traits an agent needs before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool produces, then source behavior, caps, and the research-use caveat, ending with a concrete example. It is somewhat long and the 'research summary, not a clinical report' disclaimer is stated twice, which is mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no annotations and no output schema, the description covers the return shape (ranked rows, feature attributions, caps) and the safety framing well. The only gap is the dangling reference to a `top_n` control the caller cannot actually set via the schema, which could mislead about tuning result size.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds value by explaining the source-routing semantics and that AVI scorers are Atlas-only, which reinforces the `scorers` and `source` fields. It is docked slightly because it discusses a `top_n` parameter (default 25, max 100) that does not exist in the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (a fuller per-scorer report of one variant's predicted molecular effects) and explicitly positions it against predict_variant_effect ('more rows per scorer') and against clinical-report siblings ('no pathogenicity classification'). An agent can distinguish it from the 20+ siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the automatic Atlas-vs-live routing (SNV from precomputed Atlas, indels/MNV via live inference), when the fallback occurs, and that `source` overrides it — which is real selection guidance. It stops short of an explicit 'use this instead of predict_variant_effect when you need per-scorer breadth' rule, leaving the router to infer it from the comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_allele_specific_effectsB
Predicted expression with the alternate allele against the reference allele: the RNA_SEQ scorer is that log fold change per gene and tissue, and RNA_SEQ_ACTIVE gives the expression level alongside it.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real behavioral context: it runs live model inference, declares output provenance (source: live), and disclaims clinical use. But it omits cost/latency implications of live inference, whether it is read-only/side-effect free, and any rate or permission constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is relevant but the opening sentence is awkwardly constructed and jargon-dense, requiring re-reading to parse the alt-vs-ref framing. Three short paragraphs are reasonably front-loaded with purpose, but the structure could be tighter and the scorer explanation clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does explain what the two scorer outputs represent (log fold change vs expression level) and the live-inference provenance, which covers most of what an agent needs. The main remaining gap is the absence of any sibling differentiation for a crowded name-space.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents chromosome, position, ref, alt, and tissue_type in detail (including format, 1-based hg38, and CURIE examples). The description only implicitly touches tissue filtering via 'per gene and tissue,' adding little beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation: predicting expression for the alternate allele against the reference, naming the concrete scorers (RNA_SEQ for log fold change, RNA_SEQ_ACTIVE for expression level). It is clear on the resource and output semantics, but it never distinguishes itself from sibling tools like predict_expression_impact, compare_alleles, or predict_variant_effect that sound highly overlapping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives some context: live inference, applicability to SNVs, indels and MNVs, and that results carry source: live. However, there is no explicit guidance on when to choose this over the many sibling prediction tools (e.g., predict_expression_impact or compare_alleles), leaving the choice to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_chromatin_impactA
Predicted chromatin accessibility effects of a variant (ATAC-seq and DNase-seq).
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that inference is live (score_variant with SDK-recommended scorers), that output is tagged source: live, and that results are raw model scores/quantiles with no pathogenic/benign call. It does not mention cost, latency, or failure modes of live inference, so it stops short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler; the modality and live-inference facts are front-loaded and the research-use caveat closes the description. Slightly dense but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param prediction tool with no output schema and no annotations, the description does explain what the return carries (live scores and calibrated quantiles, source tag) and sets expectations about research vs clinical use. Missing only operational details like runtime/cost and whether tissue filtering changes output shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the required allele, position and chromosome parameters are fully documented by the schema; the description adds no syntax or format detail. The optional tissue_type default behavior is only covered in the schema, not the description. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (predicted chromatin accessibility effects of a variant) and names the assay modalities (ATAC-seq and DNase-seq), which cleanly distinguishes it from sibling modality tools like predict_expression_impact, predict_splice_impact and predict_tf_binding_impact. An agent can select it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Variant classes supported (SNV, indel, MNV) and the fact that it runs live inference are stated, which implies when it applies, but there is no explicit when-to-use/ when-not guidance nor a named alternative among the many sibling predictors. Usage must be inferred from the modality name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_expression_impactA
Predicted gene expression effects of a variant: RNA-seq (log fold change per gene and tissue) and CAGE.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does fairly well: it discloses that inference is live (via score_variant with SDK-recommended scorers), that the response is marked `source: live`, and that no pathogenic/benign call is made. Missing are cost/latency implications of live inference and any failure-mode notes, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The output scope is front-loaded in the first line, followed by computational behavior and the research-use disclaimer. It is tight for the number of concepts covered, though the disclaimer sentence is somewhat verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does describe the return content (per-gene/tissue log fold change, CAGE, calibrated quantiles, source marker), and with no annotations it covers the research-only caveat. It is largely complete for a prediction tool, though the absence of return-shape or failure detail is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema, making 3 the baseline. The description's mention of 'per gene and tissue' loosely reinforces the tissue_type filter but adds no syntax or format detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('predicted gene expression effects of a variant') and pins the output modalities (RNA-seq log fold change per gene/tissue, CAGE), which lets an agent distinguish it from modality-specific siblings like predict_splice_impact or predict_tf_binding_impact. It does not explicitly contrast itself with the near-neighbor predict_variant_effect, leaving some ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the applicable variant classes (SNV, indel, MNV) and that it runs live inference, which is useful scoping. But it offers no when-to-use/when-not guidance relative to alternatives such as predict_tissue_specific or assess_pathogenicity, so routing is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_splice_impactA
Predicted splicing effects of a variant: splice sites, splice site usage and splice junctions, with the gene and junction of each.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that inference runs live via score_variant, marks the result with source: live, defines the accepted variant classes, and warns that outputs are model predictions with no pathogenic/benign call. It omits runtime cost or failure modes, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short blocks, led by the output scope, then the inference behavior, then the caveat — front-loaded and each sentence earns its place. Slightly redundant restatement of the non-clinical caveat keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by enumerating the returned elements (splice sites, usage, junctions, gene and junction per result) and flagging source: live. An agent has enough to invoke it correctly, though the shape of returned scores/quantiles is only sketched.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so position, ref, alt, chromosome, and tissue_type are fully documented in the schema already. The description adds no syntax or format guidance beyond that, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-effect and resource: predicted splicing effects of a variant, enumerating the outputs (splice sites, splice site usage, splice junctions). It is clearly distinguishable from generic siblings like predict_variant_effect or predict_tf_binding_impact without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear scope (single-nucleotide variants, indels, multi-nucleotide variants) and implies the intended use case of research prioritization rather than clinical classification, which implicitly routes away from assess_pathogenicity. However, it never explicitly names when to prefer a sibling such as predict_variant_effect for a broader effect profile.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_tf_binding_impactA
Predicted transcription factor binding effects of a variant (TF ChIP-seq), with the factor and cell type of each.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that inference is live (score_variant with recommended scorers), that the response carries `source: live`, and that results are model predictions with scores and calibrated quantiles and no pathogenic/benign call. It omits cost/latency or permission requirements, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with what the tool predicts before the mechanics and the disclaimer. Every sentence earns its place, though the boundary/limitation sentence is somewhat dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description usefully fills the gap: it names the returned contents (factor and cell type per track), the `source: live` marker, and the non-clinical nature of the scores. A 5 would require more on usage routing or failure/edge-case behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so chromosome, position, ref, alt and tissue_type are already fully documented in the schema. The description adds nothing parameter-specific (e.g., no note on tissue_type filtering behavior beyond the schema's own text), so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: predicted transcription factor binding effects of a variant, scoped to TF ChIP-seq. This modality specificity distinguishes it from siblings like predict_chromatin_impact or predict_splice_impact, though it never names a sibling explicitly to reinforce the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use statement, no condition for choosing this over the other predict_* siblings, and no exclusions. Variant-type support is mentioned ('single-nucleotide variants, indels and multi-nucleotide variants') but that is a capability, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_tissue_specificA
The strongest predicted effect of a variant in each of several tissues: the same scores, filtered to the tracks of one tissue at a time.
Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Compare the predicted effect of chr19:44908684 T>C in brain, liver and heart"
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| tissues | No | Tissue names (brain, neuron, blood, liver, heart, lung, kidney) or ontology CURIEs. Default: brain, liver, heart | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden well: it discloses that inference runs live (latency/cost implication), that results carry `source: live`, what is returned (scores and calibrated quantiles), and that no pathogenic/benign call is made. It omits auth/permission and rate-limit details, but otherwise adds meaningful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short blocks with the core purpose front-loaded and no filler paragraphs; the live-inference and research-use notes each earn their place. The opening sentence 'the same scores' lacks an antecedent, which briefly obscures what it is being compared against.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the return content (scores, calibrated quantiles, `source: live`) plus the research-only scope, and it covers supported variant classes. It is complete enough to call correctly, with only routing to sibling tools left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds modest meaning by mapping a tissue entry to 'the tracks of one tissue at a time', but it gives no extra detail on the chromosome/position/ref/alt inputs beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('predicted effect of a variant in each of several tissues') and clarifies the output shape (same scores filtered to one tissue's tracks at a time). It is distinguishable from siblings like batch_tissue_comparison and predict_variant_effect, though it never explicitly names which sibling it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied via the example ('Compare the predicted effect of chr19:44908684 T>C in brain, liver and heart'). There is no explicit statement of when to prefer this over predict_variant_effect, compare_variants, or batch_tissue_comparison, so the agent must infer routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_variant_effectA
Predicted regulatory effect of a genetic variant, per modality: the strongest tracks of each scorer (expression, transcription start, chromatin accessibility, histone marks, transcription factor binding, splicing), each with its score, calibrated quantile, gene, tissue or cell type, and assay. From the Atlas the AlphaGenome Variant Impact (AVI) score is included.
Source: a single-nucleotide variant is answered from the precomputed AlphaGenome Atlas; an indel or multi-nucleotide variant runs live inference (score_variant). Both return the same scorers in the same shape. Chosen automatically, overridable with source, and always stated in the result.
The response is a summary, never a full score matrix: ranked rows only, capped at top_n (default 25, max 100) and at 40,000 characters.
Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.
Example: "Analyze chr19:44908684 T>C with AlphaGenome"
| Name | Required | Description | Default |
|---|---|---|---|
| alt | Yes | Alternate allele (A, C, G, T; more than one base for an indel) | |
| ref | Yes | Reference allele (A, C, G, T; more than one base for an indel) | |
| source | No | Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered. | |
| scorers | No | Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves. | |
| position | Yes | Genomic position (1-based, hg38) | |
| chromosome | Yes | Chromosome (chr1-chr22, chrX, chrY) | |
| tissue_type | No | Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues. | |
| output_types | No | Optional: modalities to report (default: one scorer per modality) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses auto/live/atlas source selection and fallback, that the result states which source answered, that output is a capped summary rather than a full matrix, and that results are model predictions not clinical calls. A minor issue is that it describes a `top_n` cap (default 25, max 100) that does not appear in the input schema, which could mislead.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and then organized into scope, source behavior, response shape, and a caveat, which is efficient paragraphing. It runs a little long and includes the unreconciled `top_n` detail, but every section largely earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, no-annotation, no-output-schema tool, the description covers the response shape (ranked rows, per-modality strongest tracks, character cap), source dispatch, and the research-only caveat. What is missing is explicit sibling differentiation and reconciliation of the `top_n` cap with the actual schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces source semantics and notes that scorer names come from atlas_list_scorers and that AVI scorers are Atlas-only, but most parameter meaning is already carried by the rich schema, and the phantom `top_n` is never reconciled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (predicted regulatory effect) and resource (genetic variant) and enumerates the modalities returned (expression, TSS, chromatin, histone, TF binding, splicing) plus the AVI score. It is clear what the tool produces, but it never explicitly contrasts itself with narrower siblings like predict_splice_impact or predict_tissue_specific, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the automatic source selection and the override via `source`, plus fallback behavior, which is useful invocation guidance, and it gives a concrete example query. However, it offers no when-to-use/when-not guidance relative to the ~20 sibling prediction tools, leaving alternative selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
24 tool updates
v0.3.0- Changed
analyze_gwas_locus14 fields changed- added
Input schema / properties / chromosome / descriptionAdded value: +"Optional: locus chromosome, for the label" - added
Input schema / properties / end / descriptionAdded value: +"Optional: locus end, for the label" - added
Input schema / properties / start / descriptionAdded value: +"Optional: locus start, for the label" - added
Input schema / properties / variants / descriptionAdded value: +"Variants to score (1-100)" - added
Input schema / properties / variants / items / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / alt / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / variants / items / properties / chromosome / patternAdded value: +"^chr([1-9]|1[0-9]|2[0-2]|X|Y)$" - added
Input schema / properties / variants / items / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / variants / items / properties / position / minimumAdded value: +1 - added
Input schema / properties / variants / items / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / ref / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +} - added
Input schema / properties / variants / maxItemsAdded value: +100
- Changed
annotate_regulatory_context5 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
assess_pathogenicity6 fields changed- changed
Input schema / properties / alt / descriptionPrevious value: -"Alternate allele"New value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - changed
Input schema / properties / position / descriptionPrevious value: -"Genomic position (1-based)"New value: +"Genomic position (1-based, hg38)" - changed
Input schema / properties / ref / descriptionPrevious value: -"Reference allele"New value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / scorersAdded value: +{ + "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / sourceAdded value: +{ + "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.", + "enum": [ + "auto", + "atlas", + "live" + ], + "type": "string" +} - changed
Input schema / properties / tissue_type / descriptionPrevious value: -"Optional: disease-relevant tissue (default: brain)"New value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Added
atlas_list_scorers - Added
atlas_lookup_variant - Added
atlas_lookup_variants - Added
atlas_scan_region - Changed
batch_modality_screen12 fields changed- removed
Input schema / properties / modality / descriptionRemoved value: -"Regulatory modality to screen" - added
Input schema / properties / variants / descriptionAdded value: +"Variants to score (1-100)" - added
Input schema / properties / variants / items / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / alt / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / variants / items / properties / chromosome / patternAdded value: +"^chr([1-9]|1[0-9]|2[0-2]|X|Y)$" - added
Input schema / properties / variants / items / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / variants / items / properties / position / minimumAdded value: +1 - added
Input schema / properties / variants / items / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / ref / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +} - added
Input schema / properties / variants / maxItemsAdded value: +100
- Changed
batch_pathogenicity_filter14 fields changed- added
Input schema / properties / scorersAdded value: +{ + "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / sourceAdded value: +{ + "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.", + "enum": [ + "auto", + "atlas", + "live" + ], + "type": "string" +} - changed
Input schema / properties / threshold / descriptionPrevious value: -"Pathogenicity threshold (0-1, default: 0.5)"New value: +"Smallest absolute quantile to keep, between 0 and 1 (default: 0.99)" - added
Input schema / properties / variants / descriptionAdded value: +"Variants to score (1-100)" - added
Input schema / properties / variants / items / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / alt / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / variants / items / properties / chromosome / patternAdded value: +"^chr([1-9]|1[0-9]|2[0-2]|X|Y)$" - added
Input schema / properties / variants / items / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / variants / items / properties / position / minimumAdded value: +1 - added
Input schema / properties / variants / items / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / ref / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +} - added
Input schema / properties / variants / maxItemsAdded value: +100
- Changed
batch_score_variants10 fields changed- removed
Input schema / properties / include_interpretationRemoved value: -{ - "description": "Include detailed clinical interpretation (default: false)", - "type": "boolean" -} - added
Input schema / properties / scorersAdded value: +{ + "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.", + "items": { + "type": "string" + }, + "type": "array" +} - changed
Input schema / properties / scoring_metric / descriptionPrevious value: -"Metric to use for scoring and ranking"New value: +"What to rank by" - added
Input schema / properties / sourceAdded value: +{ + "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.", + "enum": [ + "auto", + "atlas", + "live" + ], + "type": "string" +} - changed
Input schema / properties / top_n / descriptionPrevious value: -"Number of top variants to return (default: 10, max: 100)"New value: +"Variants to return per group (default: 10, max: 100)" - changed
Input schema / properties / variants / descriptionPrevious value: -"List of variants to analyze (1-100)"New value: +"Variants to score (1-100)" - changed
Input schema / properties / variants / items / properties / alt / descriptionPrevious value: -"Alternate allele"New value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - changed
Input schema / properties / variants / items / properties / chromosome / descriptionPrevious value: -"Chromosome"New value: +"Chromosome (chr1-chr22, chrX, chrY)" - changed
Input schema / properties / variants / items / properties / position / descriptionPrevious value: -"Position"New value: +"Genomic position (1-based, hg38)" - changed
Input schema / properties / variants / items / properties / ref / descriptionPrevious value: -"Reference allele"New value: +"Reference allele (A, C, G, T; more than one base for an indel)"
- Changed
batch_tissue_comparison13 fields changed- added
Input schema / properties / tissues / descriptionAdded value: +"Tissue names or ontology CURIEs (1-10)" - removed
Input schema / properties / tissues / minItemsRemoved value: -1 - added
Input schema / properties / variants / descriptionAdded value: +"Variants to score (1-100)" - added
Input schema / properties / variants / items / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / alt / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / variants / items / properties / chromosome / patternAdded value: +"^chr([1-9]|1[0-9]|2[0-2]|X|Y)$" - added
Input schema / properties / variants / items / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / variants / items / properties / position / minimumAdded value: +1 - added
Input schema / properties / variants / items / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / ref / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +} - added
Input schema / properties / variants / maxItemsAdded value: +100
- Changed
compare_alleles6 fields changed- added
Input schema / properties / alts / descriptionAdded value: +"Alternate alleles (1-20)" - removed
Input schema / properties / alts / items / patternRemoved value: -"^[ATGCatgc]+$" - removed
Input schema / properties / alts / minItemsRemoved value: -2 - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)"
- Changed
compare_protective_risk10 fields changed- added
Input schema / properties / protective_variant / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / protective_variant / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / protective_variant / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / protective_variant / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / protective_variant / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +} - added
Input schema / properties / risk_variant / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / risk_variant / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / risk_variant / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / risk_variant / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / risk_variant / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +}
- Changed
compare_variants10 fields changed- added
Input schema / properties / variant1 / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variant1 / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / variant1 / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / variant1 / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variant1 / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +} - added
Input schema / properties / variant2 / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variant2 / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / variant2 / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / variant2 / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variant2 / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +}
- Changed
compare_variants_same_gene13 fields changed- changed
Input schema / properties / gene_name / descriptionPrevious value: -"Optional: gene name for context"New value: +"Optional: gene symbol (e.g., APOE)" - added
Input schema / properties / variants / descriptionAdded value: +"Variants to score (1-100)" - added
Input schema / properties / variants / items / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / alt / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / variants / items / properties / chromosome / patternAdded value: +"^chr([1-9]|1[0-9]|2[0-2]|X|Y)$" - added
Input schema / properties / variants / items / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / variants / items / properties / position / minimumAdded value: +1 - added
Input schema / properties / variants / items / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / variants / items / properties / ref / patternAdded value: +"^[ATGCatgc]+$" - added
Input schema / properties / variants / items / properties / variant_idAdded value: +{ + "description": "Optional: variant identifier (e.g., rs number)", + "type": "string" +} - added
Input schema / properties / variants / maxItemsAdded value: +100 - changed
Input schema / properties / variants / minItemsPrevious value: -2New value: +1
- Changed
explain_variant_impact7 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / scorersAdded value: +{ + "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / sourceAdded value: +{ + "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.", + "enum": [ + "auto", + "atlas", + "live" + ], + "type": "string" +} - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
generate_variant_report7 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / scorersAdded value: +{ + "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / sourceAdded value: +{ + "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.", + "enum": [ + "auto", + "atlas", + "live" + ], + "type": "string" +} - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
predict_allele_specific_effects5 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
predict_chromatin_impact5 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
predict_expression_impact5 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
predict_splice_impact5 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
predict_tf_binding_impact5 fields changed- added
Input schema / properties / alt / descriptionAdded value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / chromosome / descriptionAdded value: +"Chromosome (chr1-chr22, chrX, chrY)" - added
Input schema / properties / position / descriptionAdded value: +"Genomic position (1-based, hg38)" - added
Input schema / properties / ref / descriptionAdded value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / tissue_type / descriptionAdded value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
- Changed
predict_tissue_specific5 fields changed- changed
Input schema / properties / alt / descriptionPrevious value: -"Alternate allele"New value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - changed
Input schema / properties / chromosome / descriptionPrevious value: -"Chromosome"New value: +"Chromosome (chr1-chr22, chrX, chrY)" - changed
Input schema / properties / position / descriptionPrevious value: -"Genomic position"New value: +"Genomic position (1-based, hg38)" - changed
Input schema / properties / ref / descriptionPrevious value: -"Reference allele"New value: +"Reference allele (A, C, G, T; more than one base for an indel)" - changed
Input schema / properties / tissues / descriptionPrevious value: -"List of tissues to test (default: brain, liver, heart)"New value: +"Tissue names (brain, neuron, blood, liver, heart, lung, kidney) or ontology CURIEs. Default: brain, liver, heart"
- Changed
predict_variant_effect7 fields changed- changed
Input schema / properties / alt / descriptionPrevious value: -"Alternate allele (A, T, G, or C)"New value: +"Alternate allele (A, C, G, T; more than one base for an indel)" - changed
Input schema / properties / output_types / descriptionPrevious value: -"Optional: specific analyses to run (default: all)"New value: +"Optional: modalities to report (default: one scorer per modality)" - changed
Input schema / properties / position / descriptionPrevious value: -"Genomic position (1-based, positive integer)"New value: +"Genomic position (1-based, hg38)" - changed
Input schema / properties / ref / descriptionPrevious value: -"Reference allele (A, T, G, or C)"New value: +"Reference allele (A, C, G, T; more than one base for an indel)" - added
Input schema / properties / scorersAdded value: +{ + "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / sourceAdded value: +{ + "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.", + "enum": [ + "auto", + "atlas", + "live" + ], + "type": "string" +} - changed
Input schema / properties / tissue_type / descriptionPrevious value: -"Optional: tissue context (UBERON term, e.g., \"UBERON:0001157\" for brain)"New value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
20 tool updates
- First observed
analyze_gwas_locus - First observed
annotate_regulatory_context - First observed
assess_pathogenicity - First observed
batch_modality_screen - First observed
batch_pathogenicity_filter - First observed
batch_score_variants - First observed
batch_tissue_comparison - First observed
compare_alleles - First observed
compare_protective_risk - First observed
compare_variants - First observed
compare_variants_same_gene - First observed
explain_variant_impact - First observed
generate_variant_report - First observed
predict_allele_specific_effects - First observed
predict_chromatin_impact - First observed
predict_expression_impact - First observed
predict_splice_impact - First observed
predict_tf_binding_impact - First observed
predict_tissue_specific - First observed
predict_variant_effect
TDQS
Scored across 24 tools
Several tools return essentially the same 'strongest effect per modality' summary — predict_variant_effect, assess_pathogenicity, annotate_regulatory_context, generate_variant_report and explain_variant_impact overlap heavily, differing mainly in verbosity. The comparison family (compare_variants, compare_alleles, compare_variants_same_gene, compare_protective_risk, batch_tissue_comparison) and the batch family (batch_score_variants, batch_modality_screen, batch_pathogenicity_filter, atlas_lookup_variants) also blur into each other. Legacy names that the descriptions explicitly disclaim (assess_pathogenicity, compare_protective_risk, batch_pathogenicity_filter) actively mislead selection.
All names are snake_case and almost all follow a verb_noun shape (predict_*, batch_*, compare_*, atlas_*, analyze_*, annotate_*, generate_*, explain_*). The atlas_* resource prefix and the varied verb set (predict/score/assess/analyze/compare/lookup/scan) are readable conventions rather than chaos. Minor deviation: some tools lead with an action while others lead with a scope prefix.
24 tools is on the heavy side for a variant-effect predictor, and a meaningful fraction are near-duplicates of a general scoring tool. Many could be folded into predict_variant_effect with parameters (modality, comparison mode, batch) without loss. The Atlas-vs-live distinction does justify some separate entries, but not this many.
The surface covers the full variant-scoring lifecycle: single variant, batch, region scan, Atlas lookup/list, per-modality breakdowns, tissue comparison, allele comparison and plain-language/narrative output. Gaps are minor — no non-variant sequence/annotation retrieval and no export/format options — but an agent can complete realistic prioritization workflows without dead ends.
Maintenance
Related MCP Connectors
Bioinformatics MCP for genomic variant interpretation, gene-disease evidence and literature.
Cited gene, variant (rsID) and CPIC drug–gene lookups for AI agents. Read-only, no key.
Protein analysis: ESM-2/ESMC embeddings, mutation scoring, landscape scans, ESMFold structure.
Connect AI clients to biomedical data and tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to perform genomic variant analysis using OakVar, including running annotation pipelines, managing 200+ annotator modules, querying variant databases, and generating reports in various formats.MIT
- FlicenseAqualityFmaintenanceProvides interpretable variant effect predictions for 4.2 million ClinVar variants using the EVEE API. Enables searching, comparing, and analyzing genetic variants with AI-generated mechanistic interpretations and disruption profiles.618-
- AlicenseCqualityFmaintenanceProvides AI-powered access to major biological databases for GWAS and bioinformatics research. Enables natural language queries for protein, gene, variant, pathway, and drug discovery analysis.441MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI-powered protein structure prediction and variant analysis via Docker, with tools for submitting predictions, batch processing variants, and monitoring jobs.1-