Skip to main content
Glama
README.md
# gnomAD MCP Server

## Overview

This MCP server provides a programmatic interface to the Genome Aggregation Database ([gnomAD](https://gnomad.broadinstitute.org/)) API, supporting multiple API versions (**v2.1.1**, **v3.1.2**, **v4.1.0**).  
It abstracts version-specific field and schema differences, exposing a unified API for downstream tools and users.

## Status

🚧 **Under Active Development** 🚧

This project is under active development. APIs and features may change without notice.

## Supported gnomAD API Versions

- **v4.1.0** (`gnomad_r4`)
- **v3.1.2** (`gnomad_r3`)
- **v2.1.1** (`gnomad_r2_1`)

## Supported Queries by Version

The following table summarizes which queries are available for each gnomAD API version:

| Query Type                    | Description                                                      | v2  | v3  | v4  |
|-------------------------------|------------------------------------------------------------------|-----|-----|-----|
| get_gene_info                 | Retrieve gene metadata and constraint metrics (direct lookup by gene_id/gene_symbol) | āŒ  | āŒ  | āœ…  |
| get_region_info               | Retrieve variant and summary information for a genomic region    | āŒ  | āŒ  | āœ…  |
| get_variant_info              | Retrieve variant metadata and population frequency data (by variantId) | āœ…  | āœ…  | āœ…  |
| get_clinvar_variant_info      | Retrieve ClinVar variant data and clinical significance          | āœ…  | āœ…  | āœ…  |
| get_mitochondrial_variant_info| Retrieve mitochondrial variant data and population frequencies   | āŒ  | āŒ  | āœ…  |
| get_structural_variant_info   | Retrieve structural variant (SV) data and population frequencies | āœ…  | āŒ  | āœ…  |
| get_copy_number_variant_info  | Retrieve copy number variant (CNV) data and population frequencies| āŒ  | āŒ  | āœ…  |
| search_for_genes              | Search for genes by symbol or name (no direct gene_id lookup in v2/v3) | āœ…  | āœ…  | āœ…  |
| search_for_variants           | Search for variants by ID, gene, or region                      | āœ…  | āœ…  | āœ…  |
| get_str_info                  | Retrieve short tandem repeat (STR) data and population frequencies| āŒ  | āŒ  | āœ…  |
| get_all_strs                  | Retrieve all STRs in the dataset                                 | āŒ  | āŒ  | āœ…  |
| get_variant_liftover          | Retrieve liftover mapping for a variant between genomes          | āœ…  | āŒ  | āŒ  |
| get_metadata                  | Retrieve gnomAD browser metadata and API version info            | āœ…  | āœ…  | āœ…  |

- āœ… = Supported in this version
- āŒ = Not supported in this version

## Dependencies

- Python >= 3.13
- `aiohttp >= 3.11.18`
- `fastmcp >= 2.2.1`
- `gql >= 3.5.2`
- `httpx >= 0.28.1`
- `mcp[cli] >= 1.6.0`
- `nest-asyncio >= 1.6.0`
- `pytest >= 8.3.5`
- `pytest-asyncio >= 0.26.0`

## Directory Structure

```
.
ā”œā”€ā”€ gnomad/              # Main package
│   ā”œā”€ā”€ __init__.py
│   ā”œā”€ā”€ types.py         # Type definitions
│   ā”œā”€ā”€ queries/         # GraphQL query templates
│   │   ā”œā”€ā”€ v2/         # v2.1 specific queries
│   │   ā”œā”€ā”€ v3/         # v3 specific queries
│   │   └── v4/         # v4 specific queries
│   └── schemas/         # Versioned schema files
ā”œā”€ā”€ tests/               # Test code and data
│   ā”œā”€ā”€ input/          # Test input data
│   │   ā”œā”€ā”€ analyzed_schemas/  # Analyzed schema data
│   │   ā”œā”€ā”€ schema2query/     # Schema to query conversion
│   │   └── schemas/          # Raw schema files
│   ā”œā”€ā”€ output/         # Test output data
│   │   ā”œā”€ā”€ server/     # Server test outputs
│   │   ā”œā”€ā”€ v2/         # v2.1 test outputs
│   │   ā”œā”€ā”€ v3/         # v3 test outputs
│   │   └── v4/         # v4 test outputs
│   ā”œā”€ā”€ scripts/        # Test utility scripts
│   └── tests/          # Additional test modules
ā”œā”€ā”€ server.py           # FastMCP server entrypoint
ā”œā”€ā”€ pyproject.toml      # Project metadata
ā”œā”€ā”€ README.md           # This file
└── README_tests.md     # Testing documentation
```

## Setup

### Install dependencies

```bash
uv sync
```

### Activate the virtual environment

```bash
. .venv/bin/activate
```

### Test the server

```bash
uv --directory ./ run mcp dev server.py
```

### Add the MCP server to your MCP server list (Claude, Cursor, etc.)

```json
{
    "mcpServers": {
      "gnomad": {
        "command": "uv",
        "args": ["--directory", "where you cloned the repo", "run", "server.py"],
        "env": {}
      }
    }
}
```


### Run tests

Please see [README_tests.md](./README_tests.md)

## Query & API Design

- Uses the **QueryTemplateEngine** pattern to manage version-specific GraphQL query templates.
- Currently, queries are fixed; see (`./gnomad/queries`)
   - The queries were obtained using [schema_fetcher.py](gnomad/schemas/schema_fetcher.py) and [schema_analyzer.py](gnomad/schemas/schema_analyzer.py)
   - [ ] TODO: Dynamic queries
- MCP tool endpoints are documented with detailed parameter and output descriptions.


## License

This MCP server itself is licensed under the Apache License 2.0 - see the [LICENSE](LICENSE) file for details.

This project uses the gnomAD API. Please ensure you cite gnomAD when using this tool or its outputs.

## Acknowledgements

- [gnomAD](https://gnomad.broadinstitute.org/)
- [FastMCP](https://github.com/jlowin/fastmcp)

TDQS

B3.4/5.0

Scored across 12 tools

Disambiguation5/5

Each tool has a clearly distinct purpose targeting specific genomic data types (ClinVar variants, copy number variants, genes, metadata, mitochondrial variants, regions, STRs, structural variants, general variants, liftover, gene search, variant search). The descriptions clearly differentiate what each tool retrieves, with no apparent overlap in functionality.

Naming Consistency4/5

Most tools follow a consistent 'get_*_info' or 'search_for_*' pattern, but there are minor deviations: 'get_copy_number_variant_info' uses 'variantId' while 'get_variant_info' uses 'variantId' (both camelCase), and 'get_structural_variant_info' uses 'variantId' while others use 'variant_id' (snake_case). The overall pattern is clear and readable despite these inconsistencies.

Tool Count5/5

12 tools is well-scoped for a genomic database API server covering multiple data types (variants, genes, regions, etc.) across different gnomAD versions. Each tool serves a specific purpose in retrieving different genomic entities, with no obvious redundancy.

Completeness4/5

The toolset provides comprehensive retrieval capabilities for the gnomAD domain, covering all major genomic data types with both specific lookup and search functionality. The only minor gap is the version compatibility limitations noted in tool descriptions (some tools only work with specific gnomAD versions), but agents can work around this by checking metadata or using alternative tools.

Maintenance

ActivityInactive
ResponsivenessNo issues