Skip to main content
Glama
meringlab

Official STRING Database MCP Server

STRING: Perform network clustering

string_network_clustering

Clusters proteins in a STRING interaction network to identify functional modules, returning image and interactive URLs plus detailed cluster information.

Instructions

Performs network clustering on a STRING interaction network and returns a network image URL, an interactive STRING network URL, and details about each detected cluster.

Provide a table with each detected cluster’s color, STRING-derived functional description, and any returned features that distinguish it from the others.

Use the same parameters as in the network creation step to ensure consistency. If the network already contains disconnected subgraphs, the resulting number of clusters may differ from the requested value.

Inter-cluster edges are faded by default. Use inter_cluster_edge_visibility to select a different display style.

Notes:

  • For small queries (≤5 proteins), the required_score parameter is automatically lowered to 0.

  • If only a single cluster is produced, try increasing required_score, adjusting the clustering parameter, or switching to a physical network for a sparser interaction map.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
speciesNoNCBI/STRING taxonomy ID (e.g. 9606 for human, or STRG0AXXXXX for uploaded genomes).
proteinsYesOne or more protein identifiers (optionally with values). Separate entries with newline (%0d). Numeric values (e.g. expression data) can be provided after identifiers.
network_typeNoOmit for the default functional network. Its typed view can include physical and directed regulatory attributes when STRING returns them; inspect `physical` and `regulatory.directions` before claiming those edge types. Set physical for binding, complex, or co-complex questions. Set regulatory for directed regulatory relationships between proteins.
extend_networkNoAdd specified number of additional nodes to the network based on their interaction scores. Default: 0, or 10 for single-protein queries.
network_flavorNoDefaults are typed for functional networks, evidence for physical networks, and confidence for regulatory networks. Typed returns functional pairs with any physical and directed regulatory attributes that STRING reports; it does not make every pair physical or regulatory. Typed is available only for functional networks. Set evidence or confidence only when the user requests that edge display style.
required_scoreNoMinimum interaction confidence score. Omit for STRING default filtering. Set only when a threshold is requested or a broader/narrower threshold is needed.
center_node_labelsNoCenter protein labels on nodes. Set only if the user asks to center labels.
clustering_algorithmNoLeiden identifies natural communities based on network connectivity and is the default. MCL identifies densely connected subnetworks based on connectivity flow. kmeans partitions proteins into a fixed number of clusters.
clustering_parameterNoControls clustering granularity. For Leiden: resolution parameter 0.1-10.0, default 1.0; higher values produce more, smaller clusters. For MCL: inflation parameter 1.0-10.0, default 3.0. For kmeans: number of clusters, integer >=2, default 3.
hide_disconnected_nodesNoHide unconnected nodes. Set only if the user asks to hide disconnected or unconnected proteins.
inter_cluster_edge_visibilityNoHow to display edges between clusters: faded, dotted, solid, or noshow. Defaults to faded.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.13.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that required_score is auto-lowered to 0 for ≤5 proteins, that cluster count may differ from the requested value when disconnected subgraphs exist, and that inter-cluster edges default to faded. It stops short of stating auth/permission or rate-limit behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, and the notes are cleanly separated into a bulleted block. The mid-description instruction about producing a cluster table is an unusual formatting directive that slightly muddies the description's role, but overall it is well-sized with little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be re-explained, and the edge-case notes (small queries, single-cluster fallback, disconnected subgraphs) cover the main failure modes for an 11-parameter clustering tool. Only the absence of sequencing guidance relative to sibling tools keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters including clustering_algorithm and clustering_parameter semantics. The description adds only the behavioral note that required_score is auto-lowered for small queries and points at inter_cluster_edge_visibility; baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Performs network clustering on a STRING interaction network') and enumerates the outputs (image URL, interactive URL, cluster details). It implicitly distinguishes itself from the network-creation sibling by referencing 'the network creation step,' but does not name that sibling, so differentiation is contextual rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Guidance is present but indirect: it says to reuse the network-creation parameters and offers troubleshooting ('If only a single cluster is produced, try increasing required_score...'). It never states when to choose this tool over string_visual_network or string_network_link, so the agent must infer the sequencing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.