Skip to main content
Glama
taehojo
by taehojo

generate_variant_report

Create a detailed predicted molecular effect report for a variant, including AVI score feature attributions, from Atlas or live inference. Research use only.

Instructions

A fuller report of one variant's predicted molecular effects: more rows per scorer than predict_variant_effect and, from the Atlas, the AVI score with its feature attributions (AVI_SCORE_FEATURE_IMPORTANCE). It is a research summary, not a clinical report: it contains no pathogenicity classification and no recommendation.

Source: a single-nucleotide variant is answered from the precomputed AlphaGenome Atlas; an indel or multi-nucleotide variant runs live inference (score_variant). Both return the same scorers in the same shape. Chosen automatically, overridable with source, and always stated in the result.

The response is a summary, never a full score matrix: ranked rows only, capped at top_n (default 25, max 100) and at 40,000 characters.

Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.

Example: "Generate a report for chr19:44908684 T>C"

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
altYesAlternate allele (A, C, G, T; more than one base for an indel)
refYesReference allele (A, C, G, T; more than one base for an indel)
sourceNoOptional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.
scorersNoOptional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.
positionYesGenomic position (1-based, hg38)
chromosomeYesChromosome (chr1-chr22, chrX, chrY)
tissue_typeNoOptional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed7 schema fields changedv0.3.0
    • addedInput schema / properties / alt / description
      Added value: +"Alternate allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / chromosome / description
      Added value: +"Chromosome (chr1-chr22, chrX, chrY)"
    • addedInput schema / properties / position / description
      Added value: +"Genomic position (1-based, hg38)"
    • addedInput schema / properties / ref / description
      Added value: +"Reference allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / scorers
      Added value: +{
      +  "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • addedInput schema / properties / source
      Added value: +{
      +  "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.",
      +  "enum": [
      +    "auto",
      +    "atlas",
      +    "live"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / tissue_type / description
      Added value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
  2. First observed

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses that the response is a capped summary (top_n default 25, max 100; 40,000-character ceiling), that the answering source is always reported, that atlas-only mode errors rather than falls back, and that no pathogenic/benign call is made. These are exactly the traits an agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what the tool produces, then source behavior, caps, and the research-use caveat, ending with a concrete example. It is somewhat long and the 'research summary, not a clinical report' disclaimer is stated twice, which is mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with no annotations and no output schema, the description covers the return shape (ranked rows, feature attributions, caps) and the safety framing well. The only gap is the dangling reference to a `top_n` control the caller cannot actually set via the schema, which could mislead about tuning result size.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds value by explaining the source-routing semantics and that AVI scorers are Atlas-only, which reinforces the `scorers` and `source` fields. It is docked slightly because it discusses a `top_n` parameter (default 25, max 100) that does not exist in the input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (a fuller per-scorer report of one variant's predicted molecular effects) and explicitly positions it against predict_variant_effect ('more rows per scorer') and against clinical-report siblings ('no pathogenicity classification'). An agent can distinguish it from the 20+ siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the automatic Atlas-vs-live routing (SNV from precomputed Atlas, indels/MNV via live inference), when the fallback occurs, and that `source` overrides it — which is real selection guidance. It stops short of an explicit 'use this instead of predict_variant_effect when you need per-scorer breadth' rule, leaving the router to infer it from the comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.