Skip to main content
Glama
meringlab

Official STRING Database MCP Server

STRING: Functional enrichment analysis

string_enrichment

Find overrepresented biological functions for protein sets via STRING; single proteins expand to top interactors, and results report enriched terms with FDR.

Instructions

This tool retrieves functional enrichment for a set of proteins using STRING.

  • If queried with a single protein, the tool expands the query to include the protein’s 10 most likely interactors; enrichment is performed on this set, not the original single protein.

  • For two or more proteins, enrichment is performed on the exact input set.

  • When calling related tools, use the same input parameters unless otherwise specified.

  • Focus summaries on the top categories and most relevant terms for the results. Always report FDR for each claim.

  • Report FDR as a human-readable value (e.g. 2.3e-5 or 0.023).

  • IMPORTANT: Remember to suggest showing an enrichment graph for a specific category of user interest (e.g., GO, KEGG)

  • Very large responses are capped while preserving category diversity.

  • Use expand_category to return only one category with expanded term coverage and per-term gene details.

  • If a row has preferredNames_omitted: true, do not infer which proteins are in that term from the returned rows. Use string_functional_annotation with the same proteins/species and detail_for_term set to the exact term ID.

Output fields (per enriched term):

  • category: Term category (e.g., GO Process, KEGG pathway)

  • term: Enriched term (GO ID, domain, or pathway)

  • number_of_genes: Number of input genes with this term

  • number_of_genes_in_background: Number of background genes with this term

  • ncbiTaxonId: NCBI taxon ID

  • preferredNames: Canonical protein names, only when the full per-term list is short enough to show

  • proteinCount: Number of proteins matching this term

  • preferredNames_omitted: True when the gene list was omitted instead of showing a misleading partial list

  • p_value: Raw p-value

  • fdr: False Discovery Rate (B-H corrected p-value)

  • strength: Enrichment effect size, calculated as log10(observed genes / expected genes)

  • signal: Balanced enrichment-ranking metric combining the observed/expected ratio and -log(FDR)

  • description: Description of the enriched term

Response metadata:

  • input_gene_name_mapping: Only included when displayed gene lists contain submitted identifiers that differ from STRING preferred names.

  • category_summary: Total and returned term counts per category; use expand_category for categories where truncated is true or where the user wants deeper category-specific detail.

  • truncated_categories / omitted_categories: Categories with terms not shown in the current response.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
speciesNoNCBI/STRING taxon (e.g. 9606 for human, or STRG0AXXXXX). Use only when required.
proteinsYesOne or more protein identifiers, separated by %0d.
expand_categoryNoReturn only this enrichment category with expanded term coverage and a larger per-term gene-list cutoff. Use a category from metadata.category_summary.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.13.0

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and mostly succeeds: it discloses the non-obvious single-protein auto-expansion behavior, response capping that preserves category diversity, and the `preferredNames_omitted` trap with the correct workaround. It omits rate limits, auth requirements, and cost characteristics, so it stops short of exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The behavioral bullets are well front-loaded, but the description then enumerates every output field and response-metadata key even though an output schema already exists, which is largely duplicated content. The bulk is longer than needed for the incremental value it adds.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema already documents return values, the description supplies everything else an agent needs: input-set semantics, expansion side effects, truncation handling, term-detail routing, and the fallback tool for omitted gene lists. Nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: it explains how the protein count changes what `proteins` means conceptually, and clarifies `expand_category` should take a value from metadata.category_summary and yields expanded term coverage with a larger gene-list cutoff.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('retrieves functional enrichment for a set of proteins using STRING') and scopes the input semantics precisely. It also positions the tool against siblings by naming `expand_category` and `string_functional_annotation` as follow-ups, so an agent can distinguish it from adjacent enrichment tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use rules are given: a single protein triggers auto-expansion to 10 interactors, whereas two or more use the exact input set. It also states when to call `expand_category` (truncated or deeper detail) and when to switch to `string_functional_annotation` (preferredNames_omitted rows), plus the instruction to reuse the same parameters across related calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.