Skip to main content
Glama
taehojo
by taehojo

atlas_scan_region

Scan a genomic region in the AlphaGenome Atlas to rank all single-nucleotide substitutions by predicted effect without running the model. Returns top positions for prioritization.

Instructions

Scan a genomic region in the AlphaGenome Atlas: every possible single-nucleotide substitution in the interval, ranked. Answers "which positions in this region matter most?" without running the model.

Region width: at most 10,000 bp. Up to 50,000 bp only with allow_large_region=true. A scan is one API request per 32 bp under a requests-per-minute quota, so a large scan can take minutes. If the quota or the time limit stops a scan early, the partial result is returned, marked "Incomplete", with the range that was really scanned; the ranking then covers that range only.

Default scorer: AVI_SCORE (about 7 seconds for 2,000 bp). Multi-track scorers are much slower (2,000 bp: DNASE 17 s, CHIP_TF 86 s). Scorers with one row per gene or junction (RNA_SEQ, SPLICE_JUNCTIONS, ...) cannot be used for a scan: scan with AVI_SCORE, then use atlas_lookup_variant on the top variants.

The Atlas holds precomputed AlphaGenome scores for single-nucleotide substitutions on the human reference genome (hg38, chr1-22, chrX, chrY). Indels and multi-nucleotide variants are not in it; use predict_variant_effect for those.

The response is a summary, never a full score matrix: ranked rows only, capped at top_n (default 25, max 100) and at 40,000 characters.

Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.

Example: "Scan chr17:49209289-49211289 and show the 10 substitutions with the largest predicted effect"

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
endYesEnd position (greater than start; at most start + 10,000, or start + 50,000 with allow_large_region)
startYesStart position (1-based, hg38)
top_nNoRows to return (default: 25, max: 100)
scorersNoOptional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.
chromosomeYesChromosome (chr1-chr22, chrX, chrY)
allow_large_regionNoSet to true to scan more than 10,000 bp (up to 50,000 bp). Slower, and the result may be incomplete (default: false)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.3.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: 10,000 bp cap (50,000 with allow_large_region), one API request per 32 bp under a requests-per-minute quota, minutes-long runtime, partial results marked "Incomplete" with the actually-scanned range, and scorer-dependent timings (AVI ~7 s vs DNASE 17 s vs CHIP_TF 86 s for 2,000 bp). It also caps and characterizes the response and disclaims clinical use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then layers limits, timing, scorer rules, and response shape, with an example at the end. Dense and mostly waste-free, though at roughly 280 words it is longer than strictly necessary and could merge the timing and quota sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return contract itself: ranked rows only (never a full matrix), capped at top_n (default 25, max 100) and 40,000 characters, with "Incomplete" partials and their real range. Together with genome-build scope (hg38) and the research-only caveat, nothing needed to invoke or interpret the call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds real meaning beyond it, notably that scorers must come from atlas_list_scorers, that AVI scorers are Atlas-only, and that allow_large_region trades speed for possible incompleteness. It does not add much on start/end/top_n beyond what the schema already spells out.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (scan) and resource (genomic region / all SNVs in an interval) and frames the question it answers ("which positions in this region matter most?" without running the model). It explicitly distinguishes itself from siblings atlas_lookup_variant and predict_variant_effect, so an agent can route without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use and when-not: indels/MNV must go to predict_variant_effect, and multi-row scorers like RNA_SEQ cannot be scanned — instead scan with AVI_SCORE and follow up with atlas_lookup_variant on top variants. Alternatives and the conditions selecting them are named directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.