mcp-bionemo
Provides tools for NVIDIA BioNeMo biology NIMs, including RFdiffusion for protein backbone design, ProteinMPNN for amino-acid sequence design, and Boltz-2 for co-folding protein chains and optional affinity estimates. It can run against a simulator by default or route to live BioNeMo NIM endpoints.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-bionemodesign a binder against the SARS-CoV-2 spike protein and fold the complex"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-bionemo
A Model Context Protocol (MCP) server for NVIDIA BioNeMo biology NIMs. It exposes RFdiffusion, ProteinMPNN and Boltz-2 as typed MCP tools, so any MCP client (Claude Desktop, Cursor, an IDE agent, or NeMo Agent Toolkit's own MCP client) can discover and call them like any other tool. It runs against a deterministic simulator by default, so you can install and explore it with no GPU, no key and no network.
The gap this fills
BioNeMo ships as NeMo Agent Toolkit agent skills and as raw HTTP NIM endpoints. It does not ship as an MCP server. An MCP-native client therefore has no schema to discover binder-design models and cannot call them as tools. mcp-bionemo closes that gap: it wraps the NIM endpoints in tools with input schemas an LLM can read, turning protein design into an ordinary tool call. The same tools route to the real NIMs with one environment variable.
Related MCP server: mcp-anything
Quickstart
pip install -e .
# try it with the MCP inspector, or wire it into a client (below). No key needed: it runs the simulator.Point an MCP client at the server over stdio with the command mcp-bionemo. For Claude Desktop, add examples/claude_desktop_config.json to your config.
Tools
Tool | Model | What it does |
| RFdiffusion | Design a protein backbone against a target (contigs, optional hotspot residues) |
| ProteinMPNN | Design amino-acid sequences that fold to a backbone |
| Boltz-2 | Co-fold protein chains (and optional ligands), with an optional affinity estimate |
| RFdiffusion + ProteinMPNN | Backbone then sequences in one call |
| — | Report the active backend and the available tools |
A note on biology, not just plumbing: fold the binder together with the target to score binding. A binder folded alone does not measure binding, and the tool docstrings say so.
Simulated vs live
By default the server uses an in-process simulator that returns well-formed but meaningless structures and scores: it exercises the tool contracts and the orchestration, not the biology. To call the real BioNeMo NIMs, set:
BIONEMO_BACKEND=live
NGC_API_KEY=nvapi-xxxx # free key at https://build.nvidia.com
# or, for a self-hosted NIM:
BIONEMO_BASE_URL=https://health.api.nvidia.com/v1Hosted NIMs return HTTP 202 for long jobs; the client polls the status endpoint until the result is ready.
Development
pip install -e ".[dev]"
ruff check .
pytest # exercises every tool against the simulatorLicense
Apache-2.0. This project calls NVIDIA BioNeMo NIMs but is not affiliated with or endorsed by NVIDIA.
Available Tools
5 toolsdesign_backboneB
Design a protein backbone with RFdiffusion.
Args: input_pdb: the target structure as PDB text (the scaffold/target to design against). contigs: RFdiffusion contig string describing what to generate, e.g. "A1-100/0 50-60". hotspot_res: optional target residues the binder should contact, e.g. ["A59", "A83"].
Returns a dict with output_pdb (the generated backbone as PDB text).
| Name | Required | Description | Default |
|---|---|---|---|
| contigs | Yes | ||
| input_pdb | Yes | ||
| hotspot_res | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the return shape ('output_pdb' as PDB text) and documents all inputs, but says nothing about runtime/GPU cost, determinism or seeding, or failure modes (e.g. invalid contig strings) — notable omissions for a diffusion-model tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Structured as a purpose line followed by Args and Returns sections, with the core purpose front-loaded. No filler sentences; the argument lines are dense but each earns its place. Trailing whitespace/minor formatting is the only blemish.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so explaining return values is optional (though harmless here). Inputs and purpose are well covered for a 3-parameter tool, but the absence of any guidance on pipeline position relative to the sibling design/fold tools and the absence of runtime or determinism context leave it only minimally adequate for a specialized diffusion tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: all three parameters are given meaning plus concrete examples ('A1-100/0 50-60' for contigs, ['A59','A83'] for hotspot_res). It stops short of explaining contig mini-language rules or expected PDB formatting, but the added value over the bare schema is substantial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Design a protein backbone') and names the engine (RFdiffusion), which is concrete. It does not explicitly distinguish itself from siblings like design_sequences, design_binder, or fold_complex, so an agent must infer the boundary from 'backbone' vs. 'sequences'/'binder'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the arguments mean but never states when to choose this tool over design_binder or design_sequences, nor any prerequisites (e.g. that this is the first stage of a design pipeline). Usage is only weakly implied by the hotspot_res description mentioning a 'binder'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_binderA
Convenience pipeline: RFdiffusion backbone, then ProteinMPNN sequences for it, in one call.
Returns a dict with backbone_pdb and designs (FASTA) plus scores. Fold the designs against
the target with fold_complex to score binding.
| Name | Required | Description | Default |
|---|---|---|---|
| contigs | Yes | ||
| target_pdb | Yes | ||
| hotspot_res | No | ||
| num_sequences | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses the underlying models (RFdiffusion, ProteinMPNN) and the return keys (backbone_pdb, designs, scores), but omits operational traits like GPU cost, expected runtime, or whether the call is expensive/long-running — important for a diffusion pipeline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with zero filler, front-loaded with the core pipeline composition, then return shape, then the recommended next step. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained (though the description does list keys). For a multi-parameter pipeline tool with 0% schema description coverage, the description leaves the agent without guidance on the two required parameters, which is a meaningful gap despite adequate workflow framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four parameters (contigs, target_pdb, hotspot_res, num_sequences), and the description explains none of them. There is only an oblique reference to 'the target', which does not clarify contig syntax, hotspot residue format, or the meaning of num_sequences.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('design_binder' = design a binder) and explains that it is a pipeline combining RFdiffusion backbone generation with ProteinMPNN sequence design 'in one call'. This implicitly differentiates it from the design_backbone and design_sequences siblings, though it never names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: the pipeline nature implies you use this instead of chaining the two single-step tools. It also explicitly names the follow-up tool and why ('Fold the designs against the target with fold_complex to score binding'), which routes the agent through the workflow. No explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_sequencesB
Design amino-acid sequences that fold to a given backbone with ProteinMPNN.
Args: input_pdb: the backbone (PDB text) to design sequences for. num_sequences: how many candidate sequences to return.
Returns a dict with mfasta (FASTA of designs) and scores (lower is better).
| Name | Required | Description | Default |
|---|---|---|---|
| input_pdb | Yes | ||
| num_sequences | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it partially fulfills it by disclosing the underlying model (ProteinMPNN), the return payload (mfasta and scores), and the score direction (lower is better). It is silent on cost, determinism, input size limits, and whether designs are sampled stochastically, which are relevant behavioral traits for a generative design tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The summary sentence is front-loaded and the Args/Returns blocks are compact and scannable. Slightly verbose in listing parameters that also appear in the schema, but no sentence is wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output semantics are covered by both the description and the output schema, and both inputs are explained. However, for a generative design tool with no annotations, the description should say more about resource cost, determinism, or constraints on the backbone input to fully guide an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: input_pdb is defined as PDB text of the backbone and num_sequences as the number of candidates. This meaningfully clarifies both parameters, though it omits the default of 8.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Design) and resource (amino-acid sequences), and pins down the mechanism (ProteinMPNN) plus the input requirement (a given backbone). This clearly differs from the sibling design_backbone or fold_complex, but it never names those siblings, so differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what the tool does but gives no when-to-use guidance, no prerequisites, and no comparison to siblings like design_binder or design_backbone. An agent must infer that this is the sequence-design step rather than getting explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fold_complexA
Co-fold one or more protein chains (optionally with small-molecule ligands) using Boltz-2.
Args: protein_sequences: one amino-acid sequence per protein chain. Give the binder AND the target together to score the complex; a binder folded alone does not measure binding. ligand_smiles: optional list of SMILES strings for small-molecule ligands. predict_affinity: if true and a ligand is given, also return a binding-affinity estimate.
Returns a dict with structures, confidence_scores, and (optionally) affinities.
| Name | Required | Description | Default |
|---|---|---|---|
| ligand_smiles | No | ||
| predict_affinity | No | ||
| protein_sequences | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden: it discloses the underlying method (Boltz-2), the conditional gating of affinity output, and the top-level return keys. It omits operational traits an agent would care about, such as runtime/compute cost and any limits on chain or ligand count.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in one sentence, then uses a compact args block; the Returns line is short. The args block largely restates schema fields, so a bit of the space is not adding new information, but nothing is verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter folding tool with a machine-readable output schema, the definition covers purpose, parameter meaning and output keys adequately. It stops short of describing practical constraints (input size limits, expected runtime, whether Boltz-2 runs locally or needs external setup) that would matter for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates by defining each of the three parameters semantically: one amino-acid sequence per chain for protein_sequences, optional SMILES strings for ligands, and the ligand-dependent behavior of predict_affinity. Minor gap: how multiple SMILES map to chains is not explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (co-fold), resource (protein chains, optionally with ligands) and names the engine (Boltz-2), which clearly separates it from the design_* siblings that generate sequences or backbones. An agent can tell from the first sentence that this tool scores/evaluates existing sequences rather than designing new ones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage condition: give binder AND target together to score a complex, because a binder folded alone does not measure binding, and states that affinity is only returned when a ligand is supplied and predict_affinity is true. It does not name or exclude the design_* siblings explicitly, so routing to alternatives is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infoA
Return which backend is active (simulated or live) and the tools this server exposes.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, but it does disclose the key behavioral fact an agent cares about: which backend mode is active (simulated vs live). It is a zero-parameter read with no stated side effects, which is inferable, but it omits any note on stability of the returned tool list or auth/rate considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It names both things the tool returns without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't enumerate return fields, and it correctly summarizes the two return categories. For a zero-param info tool this is nearly complete; only the lack of any usage framing keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; no parameter information could add value here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and two concrete resources: the active backend mode (simulated or live) and the list of exposed tools. This is self-evidently distinct from the sibling design_* tools, though the description does not name or reference those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are given. An agent can infer this is a diagnostic/introspection call, but the description never says so or contrasts it with the design tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
design_backbone - First observed
design_binder - First observed
design_sequences - First observed
fold_complex - First observed
info
TDQS
Scored across 5 tools
The tools mostly target distinct stages of the protein-design workflow: info, backbone generation, sequence design, and complex folding. The only overlap is design_binder, which is a convenience wrapper around design_backbone plus design_sequences, but its description makes that relationship clear.
Tool names are consistently snake_case and mostly follow an action-oriented pattern: design_backbone, design_sequences, design_binder, fold_complex. The lone outlier is info, which is a noun-only metadata tool, but it does not break readability.
Five tools is well-scoped for a focused protein/binder design server. Each tool has a clear role in the pipeline, and there is no redundant bloat beyond the intentional convenience wrapper.
The surface covers the core workflow: backbone design, sequence design, complex folding, optional affinity estimation, and a combined pipeline. Minor gaps remain around ranking/filtering designs and direct protein-protein affinity scoring, but agents can work around these using the returned confidence scores.
Maintenance
Related MCP Connectors
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
MCP server for progressive tool usage at any scale (see https://klavis.ai)
Connect AI clients to biomedical data and tools.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to discover and execute tools via a secure MCP server with JWT authentication, RBAC, rate limiting, and audit logging.1MIT
- AlicenseAqualityBmaintenanceEnables models to discover, inspect, and call tools from thousands of MCP servers on the fly through a single gateway, without pre-configuring each server.547 npm3MIT
- AlicenseAqualityAmaintenanceEnables AI harnesses to connect to a single MCP endpoint that routes to multiple downstream MCP servers, discovering and executing capabilities on demand while keeping tool schemas out of context.470 npmApache 2.0
- AlicenseNot gradedqualityAmaintenanceEnables AI tools to uniformly discover, inspect, and call tools, prompts, and resources from multiple upstream MCP servers through a small set of fixed MCP tools, over stdio or HTTP.2MIT