Skip to main content
Glama
taehojo
by taehojo

predict_tf_binding_impact

Predict how a DNA variant alters transcription factor binding, with the affected factor and cell type, to prioritize variants for genomic research.

Instructions

Predicted transcription factor binding effects of a variant (TF ChIP-seq), with the factor and cell type of each.

Runs live inference (score_variant with the SDK's recommended variant scorers); works for single-nucleotide variants, indels and multi-nucleotide variants. The result states source: live.

Results are AlphaGenome model predictions for research prioritization, not clinical classifications: scores and calibrated quantiles are reported as returned, and no pathogenic/benign call is made.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
altYesAlternate allele (A, C, G, T; more than one base for an indel)
refYesReference allele (A, C, G, T; more than one base for an indel)
positionYesGenomic position (1-based, hg38)
chromosomeYesChromosome (chr1-chr22, chrX, chrY)
tissue_typeNoOptional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv0.3.0
    • addedInput schema / properties / alt / description
      Added value: +"Alternate allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / chromosome / description
      Added value: +"Chromosome (chr1-chr22, chrX, chrY)"
    • addedInput schema / properties / position / description
      Added value: +"Genomic position (1-based, hg38)"
    • addedInput schema / properties / ref / description
      Added value: +"Reference allele (A, C, G, T; more than one base for an indel)"
    • addedInput schema / properties / tissue_type / description
      Added value: +"Optional: keep only the tracks of one tissue or cell type. A name (brain, neuron, blood, liver, heart, lung, kidney) or an ontology CURIE (e.g., UBERON:0000955, CL:0000540). Default: all tissues."
  2. First observed

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that inference is live (score_variant with recommended scorers), that the response carries `source: live`, and that results are model predictions with scores and calibrated quantiles and no pathogenic/benign call. It omits cost/latency or permission requirements, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with what the tool predicts before the mechanics and the disclaimer. Every sentence earns its place, though the boundary/limitation sentence is somewhat dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description usefully fills the gap: it names the returned contents (factor and cell type per track), the `source: live` marker, and the non-clinical nature of the scores. A 5 would require more on usage routing or failure/edge-case behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so chromosome, position, ref, alt and tissue_type are already fully documented in the schema. The description adds nothing parameter-specific (e.g., no note on tissue_type filtering behavior beyond the schema's own text), so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: predicted transcription factor binding effects of a variant, scoped to TF ChIP-seq. This modality specificity distinguishes it from siblings like predict_chromatin_impact or predict_splice_impact, though it never names a sibling explicitly to reinforce the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use statement, no condition for choosing this over the other predict_* siblings, and no exclusions. Variant-type support is mentioned ('single-nucleotide variants, indels and multi-nucleotide variants') but that is a capability, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.