Skip to main content
Glama
taehojo
by taehojo

analyze_gwas_locus

Rank variants in a GWAS locus by predicted effect across modalities to prioritize follow-up candidates. Scores SNVs, indels, and multi-nucleotide variants using live AlphaGenome predictions.

Instructions

Rank the variants of a locus by predicted effect (largest absolute quantile across modalities), to prioritize candidates for follow-up. For single-nucleotide variants only, atlas_lookup_variants or atlas_scan_region is faster and adds the AVI score.

Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.

Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
endNoOptional: locus end, for the label
startNoOptional: locus start, for the label
variantsYesVariants to score (1-100)
chromosomeNoOptional: locus chromosome, for the label

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed14 schema fields changedv0.3.0
    • addedInput schema / properties / chromosome / description
      Added value: +"Optional: locus chromosome, for the label"
    • addedInput schema / properties / end / description
      Added value: +"Optional: locus end, for the label"
    • addedInput schema / properties / start / description
      Added value: +"Optional: locus start, for the label"
    • addedInput schema / properties / variants / description
      Added value: +"Variants to score (1-100)"
    • addedInput schema / properties / variants / items / properties / alt / description
      Added value: +"Alternate allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / variants / items / properties / alt / pattern
      Added value: +"^[ATGCatgc]+$"
    • addedInput schema / properties / variants / items / properties / chromosome / description
      Added value: +"Chromosome (chr1-chr22, chrX, chrY)"
    • addedInput schema / properties / variants / items / properties / chromosome / pattern
      Added value: +"^chr([1-9]|1[0-9]|2[0-2]|X|Y)$"
    • addedInput schema / properties / variants / items / properties / position / description
      Added value: +"Genomic position (1-based, hg38)"
    • addedInput schema / properties / variants / items / properties / position / minimum
      Added value: +1
    • addedInput schema / properties / variants / items / properties / ref / description
      Added value: +"Reference allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / variants / items / properties / ref / pattern
      Added value: +"^[ATGCatgc]+$"
    • addedInput schema / properties / variants / items / properties / variant_id
      Added value: +{
      +  "description": "Optional: variant identifier (e.g., rs number)",
      +  "type": "string"
      +}
    • addedInput schema / properties / variants / maxItems
      Added value: +100
  2. First observed

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that inference runs live via score_variant with recommended scorers, that the result is tagged source: live, that it covers SNVs/indels/MNVs, and that outputs are research predictions with no pathogenic/benign call. It does not mention cost, rate limits, or latency for a 1-100 variant live call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short paragraphs, purpose and alternative front-loaded before the mechanism and disclaimer. Mostly efficient, though the SNV phrasing in the first paragraph and the modality coverage in the second overlap enough to introduce mild ambiguity rather than pure redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations exist, so the description must supply behavioral context, and it covers the live-inference source flag, modality scope, and non-clinical disclaimer. It stops short of describing the shape of the ranking output or how many results are returned, which would fully complete it for a 1-100 variant tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so start/end/chromosome/variants are already well documented including the 1-100 bound. The description adds no additional semantics (e.g., interpretation of the locus label fields) beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (rank) and resource (variants of a locus) plus the ranking criterion (largest absolute quantile across modalities) and the goal (prioritize follow-up candidates). An agent can distinguish it from sibling scorers like predict_variant_effect or batch_score_variants from this alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes SNV-only inputs to atlas_lookup_variants or atlas_scan_region, naming the condition (SNV-only) and the benefit (faster, adds AVI). However the phrasing 'For single-nucleotide variants only' sits awkwardly beside the later claim that the tool handles SNVs, indels and MNVs, leaving the exact when-not boundary slightly muddy.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.