Skip to main content
Glama
taehojo
by taehojo

batch_pathogenicity_filter

Filter and rank variants whose predicted effect reaches a set quantile threshold, using AlphaGenome scores for research prioritization.

Instructions

Keep the variants whose predicted effect reaches a threshold, ranked. The tool name is kept for compatibility: it filters on predicted effect size, NOT on pathogenicity, and classifies nothing.

threshold is the smallest absolute calibrated quantile (0 to 1) a variant must reach to be kept (default 0.99): the AVI score's quantile for variants answered from the Atlas, the largest quantile across modalities for live inference. Variants are routed per variant and reported per source; groups from different sources are not comparable.

Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.

Example: "Which of these variants have a predicted effect above the 99.9th percentile?"

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
sourceNoOptional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.
scorersNoOptional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.
variantsYesVariants to score (1-100)
thresholdNoSmallest absolute quantile to keep, between 0 and 1 (default: 0.99)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed14 schema fields changedv0.3.0
    • addedInput schema / properties / scorers
      Added value: +{
      +  "description": "Optional: scorer names to use instead of the defaults. Names come from atlas_list_scorers and are the same for both sources, except the AVI scorers, which the Atlas alone serves.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • addedInput schema / properties / source
      Added value: +{
      +  "description": "Optional: where the answer comes from (default: auto). auto = the precomputed AlphaGenome Atlas for single-nucleotide substitutions, live inference for everything else (indels, multi-nucleotide variants); falls back to live only when the Atlas does not hold the variant. atlas = Atlas only, errors instead of falling back. live = always run the model. The result always states which source answered.",
      +  "enum": [
      +    "auto",
      +    "atlas",
      +    "live"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / threshold / description
      Previous value: -"Pathogenicity threshold (0-1, default: 0.5)"New value: +"Smallest absolute quantile to keep, between 0 and 1 (default: 0.99)"
    • addedInput schema / properties / variants / description
      Added value: +"Variants to score (1-100)"
    • addedInput schema / properties / variants / items / properties / alt / description
      Added value: +"Alternate allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / variants / items / properties / alt / pattern
      Added value: +"^[ATGCatgc]+$"
    • addedInput schema / properties / variants / items / properties / chromosome / description
      Added value: +"Chromosome (chr1-chr22, chrX, chrY)"
    • addedInput schema / properties / variants / items / properties / chromosome / pattern
      Added value: +"^chr([1-9]|1[0-9]|2[0-2]|X|Y)$"
    • addedInput schema / properties / variants / items / properties / position / description
      Added value: +"Genomic position (1-based, hg38)"
    • addedInput schema / properties / variants / items / properties / position / minimum
      Added value: +1
    • addedInput schema / properties / variants / items / properties / ref / description
      Added value: +"Reference allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / variants / items / properties / ref / pattern
      Added value: +"^[ATGCatgc]+$"
    • addedInput schema / properties / variants / items / properties / variant_id
      Added value: +{
      +  "description": "Optional: variant identifier (e.g., rs number)",
      +  "type": "string"
      +}
    • addedInput schema / properties / variants / maxItems
      Added value: +100
  2. First observed

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does real work: it explains that variants route per variant, results report per source, and groups from different sources are not comparable. It also discloses that no pathogenic/benign call is made and that scores/quantiles are returned as-is, though it does not cover error or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the key name-correction caveat, then threshold semantics, then the disclaimer and an example. Mostly efficient, though the 'not clinical classifications' disclaimer is restated twice ('no pathogenic/benign call is made'), which is slight redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does supply return context: results are ranked, reported per source, and scores and calibrated quantiles are returned as-is. Combined with the schema-documented source routing, an agent has enough to call and interpret the tool, though pagination/volume of returns is not addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3, but the description adds genuine meaning to `threshold` beyond the schema: the smallest absolute calibrated quantile for Atlas-answered variants versus the largest quantile across modalities for live inference. That source-dependent semantics is not available from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: it keeps variants whose predicted effect reaches a threshold, ranked. It explicitly corrects its misleading name ('filters on predicted effect size, NOT on pathogenicity, and classifies nothing'), which is exactly the disambiguation an agent needs against siblings like assess_pathogenicity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It makes the use context clear (effect-size prioritization, not clinical classification) and gives a concrete triggering question as an example. It stops short of naming an alternative tool for the pathogenicity-classification use case it rules out, leaving that route implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.