Skip to main content
Glama

Batch Search and Download

encode_batch_download
Idempotent

Search ENCODE files by filters (format, assay, organism) and batch download matches to a local folder, with dry-run preview and MD5 verification.

Instructions

Search for files and download them all in batch.

First searches for files matching the criteria, then downloads them. By default runs in dry_run mode to preview what would be downloaded. Set dry_run=False to actually download.

WHEN TO USE: Use for searching and downloading files in one step. Always use dry_run=True first to preview. For specific file accessions, use encode_download_files. RELATED TOOLS: encode_download_files, encode_search_files

Examples:

  • Download all BED files from human pancreas ChIP-seq: file_format="bed", assay_title="Histone ChIP-seq", organ="pancreas", download_dir="/data/encode", dry_run=False

  • Preview FASTQ downloads for mouse brain RNA-seq: file_format="fastq", assay_title="total RNA-seq", organ="brain", organism="Mus musculus", download_dir="/data/encode"

  • Download IDR peaks for H3K27me3 in GRCh38: output_type="IDR thresholded peaks", target="H3K27me3", assembly="GRCh38", download_dir="/data/encode", dry_run=False

Args: download_dir: Local directory to save files file_format: File format filter ("fastq", "bam", "bed", "bigWig", etc.) output_type: Output type filter ("reads", "peaks", "signal", etc.) output_category: Output category ("raw data", "alignment", "annotation", etc.) assembly: Genome assembly ("GRCh38", "mm10", etc.) assay_title: Assay type ("Histone ChIP-seq", "ATAC-seq", "total RNA-seq", etc.) organism: Organism (default: "Homo sapiens") organ: Organ/tissue ("pancreas", "brain", "liver", etc.) biosample_type: Biosample type ("tissue", "cell line", "primary cell", etc.) target: ChIP/CUT&RUN target ("H3K27me3", "CTCF", etc.) preferred_default: If True, only download default/recommended files organize_by: File organization ("flat", "experiment", "format", "experiment_format") verify_md5: Verify downloads with MD5 checksums (default True) limit: Max files to download (default 100, safety limit) dry_run: If True (default), only preview what would be downloaded. Set False to download. offset: Skip the first N matching files; pass the next_offset of the previous reply to continue a search that has more files than limit

Returns: JSON with download preview (dry_run=True) or download results (dry_run=False).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNo
organNo
offsetNo
targetNo
dry_runNo
assemblyNo
organismNoHomo sapiens
verify_md5No
assay_titleNo
file_formatNo
organize_byNoexperiment
output_typeNo
download_dirYes
biosample_typeNo
output_categoryNo
preferred_defaultNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.3.4
    • addedInput schema / properties / offset
      Added value: +{
      +  "default": 0,
      +  "title": "Offset",
      +  "type": "integer"
      +}
  2. First observedv0.3.0-beta.1

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful operational detail beyond annotations: default dry_run behavior, the search-then-download sequence, MD5 verification, a safety limit of 100 files, and pagination via offset/next_offset. It does not mention auth or existing-file overwrite behavior, but the idempotent/non-destructive annotation profile makes those less critical.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but well-organized into WHEN TO USE, RELATED TOOLS, examples, Args, and Returns. Dry-run guidance is repeated a few times, and the examples are somewhat rich, but for a 16-parameter tool the structure keeps it scannable and each major section carries useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 16-parameter interface, open-world search semantics, and a multi-step download workflow, the description covers the full calling context: two-phase behavior, dry-run safety, filtering dimensions, pagination, organization options, and return type. The output schema covers detailed return fields, so nothing operationally essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet every one of the 16 parameters is individually documented in the Args list with meanings, defaults, allowed examples, and usage notes. For parameters like offset, it even explains how to use next_offset to continue interrupted searches, fully compensating for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Search for files and download them all in batch,' and distinguishes this tool from encode_download_files for specific accessions. An agent can immediately tell this is the batch search-and-download path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The WHEN TO USE section explicitly says to use this for combining search and download in one step, instructs dry_run=True first, and identifies encode_download_files as the alternative for specific accessions. This gives concrete decision criteria for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.