Heuris-BioMCP
Integrates with NVIDIA NIM to run AI models such as Boltz-2 for protein structure prediction and Evo2 for DNA generation and scoring.
Exposes metrics in Prometheus format for monitoring server performance and usage.
Allows searching PubMed literature with MeSH terms, Boolean syntax, and retrieval of abstracts and metadata.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Heuris-BioMCPFind clinical trials for EGFR"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🧬 Heuris-BioMCP — Bioinformatics Model Context Protocol Server
Strategic Model Context Protocol server for life sciences.
Connect ChatGPT, Claude, and other MCP clients to a curated biology tool surface built for research, translational workflows, and production review.
🚀 Quick Start • 🔧 Tools • 📊 Databases • 💡 Examples • 🤝 Contributing
Live Demo
Try Heuris-BioMCP without installing — connect to our live server:
https://heuris-biomcp.onrender.com/mcpIf hosted auth is enabled, Claude-compatible clients will complete the OAuth redirect flow automatically. Non-interactive clients can also use bearer API keys.
Connect to Live Server
For Claude, use Customize > Connectors and enter:
https://heuris-biomcp.onrender.com/mcpFor generic remote MCP clients that accept URL-based server definitions:
{
"mcpServers": {
"heuris-biomcp": {
"url": "https://heuris-biomcp.onrender.com/mcp"
}
}
}Note:
https://heuris-biomcp.onrender.com/mcpis the recommended remote MCP endpoint for modern clients. The legacy SSE endpoint remains available athttps://heuris-biomcp.onrender.com/sse.
Related MCP server: OrigeneMCP
See Heuris-BioMCP in Action

Quick Demo Video
Watch how to connect Heuris-BioMCP to Claude Desktop:
Tip: Coming soon - video walkthrough of connecting Heuris-BioMCP and running your first query!
What is Heuris-BioMCP?
Heuris-BioMCP bridges MCP clients and high-value life-science data sources through a curated public tool surface. The server exposes the workflows that matter most for product review and research use, while lower-level helper modules remain internal for composition and planner logic.
You -> "What drugs target EGFR and what clinical trials are recruiting?"
ChatGPT + Heuris-BioMCP -> Queries ChEMBL + ClinicalTrials.gov -> Structured answerWhat's New in v2.3
Curated the public MCP surface from a broad 71-tool registry down to a strategy-driven set of 32 review-friendly tools
Merged operational suites into workflow tools:
find_protein,pathway_analysis,crispr_analysis,drug_safety,variant_analysis, andsessionAdded new translational tools:
drug_interaction_checker,protein_binding_pocket,biomarker_panel_design,pharmacogenomics_report,protein_family_analysis,network_enrichment,rnaseq_deconvolution,structural_similarity,rare_disease_diagnosis, andgenome_browser_snapshotRemoved low-signal or niche tools from the public MCP registry while keeping lower-level code available internally
Preserved hosted HTTP/SSE deployment and operational endpoints for production-style MCP review
Added optional hosted OAuth 2.1 PKCE and API-key auth for authenticated remote connectors
Added Prometheus metrics, streamed progress chunks for slow tools, persistent gene/disease/watch MCP resources, and literature watch workflows
Tools (32 curated public surface)
Core Research
Tool | Description |
| PubMed literature search with MeSH, Boolean syntax, abstracts, and metadata |
| NCBI Gene summary with aliases, locus, and functional context |
| NCBI BLAST sequence alignment |
| Full UniProt Swiss-Prot protein record |
| Unified UniProt plus PDB protein discovery workflow |
| AlphaFold structure metadata and confidence summary |
| Merged KEGG plus Reactome pathway workflow |
| ChEMBL drug-target evidence for a gene |
| Open Targets translational gene-disease evidence |
| ClinicalTrials.gov recruiting-trial search |
| Integrated multi-database flagship gene report |
AI And Engineering Workflows
Tool | Description |
| Boltz-2 structure workflow with optional protein-ligand mode |
| Evo2 generation or WT-vs-variant scoring workflow |
| Merged CRISPR design, scoring, off-target, base-edit, and repair workflow |
| Merged FDA safety workflow for events, signals, labels, and comparisons |
| Merged ACMG, gnomAD, ClinVar, splice, and integrated variant reporting |
| Merged entity, graph, export, and adaptive planning workflow |
High-Value Translational Tools
Tool | Description |
| Drug repurposing workflow over literature, trials, and target evidence |
| Cross-database biological claim verification |
| Cancer mutation frequency search |
| GWAS trait-association search |
| FDA label-based interaction screening |
| Candidate binding-site summary from annotated protein features |
| Disease-focused biomarker panel drafting |
| CPIC-style pharmacogenomics summary with PGx evidence |
| Protein family and domain context |
| Gene-set pathway and interaction-hub enrichment summary |
| Marker-based bulk RNA-seq deconvolution |
| PubChem-based chemical structural similarity search |
| Phenotype normalization plus OMIM-oriented rare-disease differential support |
| Browser-ready locus context for genes and genomic intervals |
Public Surface Policy
The MCP registry is intentionally curated.
Lower-level legacy implementations still exist in the package for internal orchestration and testing.
Reviewers should evaluate the exposed MCP surface, not the hidden implementation inventory.
Databases & AI Models
Source | Domain | URL |
Traditional Databases | ||
NCBI PubMed | Literature | |
NCBI Gene | Genomics | |
NCBI BLAST | Sequence Alignment | |
NCBI GEO | Gene Expression | |
UniProt Swiss-Prot | Proteomics | |
AlphaFold DB | Protein Structure | |
RCSB PDB | Protein Structure | |
KEGG | Pathways | |
Reactome | Pathways | |
ChEMBL | Drug Discovery | |
Open Targets | Gene-Disease | |
Ensembl | Genomics | |
ClinicalTrials.gov | Clinical | |
Human Cell Atlas | Single-Cell | |
OpenNeuro | Neuroimaging | |
NeuroVault | Neuroimaging | |
v2 Extended Databases | ||
OMIM | Genetic Diseases | |
STRING | Protein Interactions | |
GTEx | Expression Atlas | |
cBioPortal | Cancer Genomics | |
GWAS Catalog | Trait Associations | |
DisGeNET | Disease-Gene | |
PharmGKB | Pharmacogenomics | |
v2.2 Tier 2 Databases | ||
BioGRID | Protein Interactions | |
Orphanet | Rare Diseases | |
GDC / TCGA | Tumor Genomics | |
CellMarker | Cell Type Markers | |
ENCODE | Regulatory Elements | |
MetaboLights | Metabolomics | |
UCSC Genome Browser | Splice Isoforms | |
Safety, Variant & Innovation Sources | ||
OpenFDA / FAERS | Drug Safety | |
DailyMed | Drug Labels | |
ClinVar | Clinical Variants | |
gnomAD | Population Variation | |
bioRxiv / medRxiv | Preprints | |
InterPro | Protein Domains | |
COSMIC | Cancer Mutations | |
AI Models (NVIDIA NIM) | ||
MIT Boltz-2 | Structure Prediction | |
Arc Evo2-40B | DNA Generation |
Quick Start
Option 1: Use Live Demo (No Installation)
Use Claude's Customize > Connectors flow and enter:
https://heuris-biomcp.onrender.com/mcpIf you are connecting with an older MCP client that still expects SSE, use:
https://heuris-biomcp.onrender.com/sseOption 2: Deploy Your Own
Deploy to Render with one click:
Or manually:
Fork this repository
Create a new Web Service on Render
Connect your fork
Set build command:
pip install -r requirements.txt && pip install -e .Set start command:
BIOMCP_TRANSPORT=http BIOMCP_HTTP_PORT=$PORT python -m biomcp
Hosted Deployment Limitations
Session snapshots saved through the
sessiontool are only durable ifBIOMCP_SESSION_STORE_DIRpoints to persistent storage.On Render free tier, the default local directory uses ephemeral disk and will be wiped on restart, redeploy, or scale-to-zero wake-up.
If you need persistent saved sessions, set
BIOMCP_SESSION_STORE_DIRto a mounted persistent path or move session persistence behind an external store before relying on cross-session restore.
Hosted Auth and Connector Setup
For Anthropic-style hosted connectors, enable
BIOMCP_AUTH_ENABLED=1.OAuth 2.1 PKCE is exposed at
/.well-known/oauth-authorization-server,/oauth/authorize,/oauth/token, and/oauth/register.For machine-to-machine clients, set
BIOMCP_API_KEYSand send eitherAuthorization: Bearer <key>orX-API-Key: <key>.Per-key limits are controlled with
BIOMCP_API_KEY_RATE_LIMIT_REQUESTSandBIOMCP_API_KEY_RATE_LIMIT_WINDOW_SECONDS.
Persistent Resources and Literature Watches
biomcp://gene/{SYMBOL}returns a curated gene context resource with gene, protein, pathway, disease, and drug-target context.biomcp://disease/{URL-ENCODED-NAME}returns disease literature plus session-graph context.session(action="watch")registers a PubMed + bioRxiv watch and exposesbiomcp://watch/{TOPIC}as a reusable resource.Saved sessions and watches are only persistent if the backing session-store directory is durable.
Privacy, Support, and Data Handling
Privacy policy: PRIVACY.md
Support channel: SUPPORT.md
Data handling notes: DATA_PROCESSING.md
Security reporting: SECURITY.md
Option 3: Local Installation
Prerequisites
Python 3.11+
Claude Desktop or any MCP-compatible client
(Optional) NCBI API key for higher rate limits
(Optional) NVIDIA API keys for AI tools
Installation
# Clone the repository
git clone https://github.com/SachinGawande2003/Heuris-BioMCP.git
cd Heuris-BioMCP
# Install (standard)
pip install -e .
# Install with neuroimaging support
pip install -e ".[neuroimaging]"
# Install with dev dependencies (recommended)
pip install -e ".[dev]"Configure Claude Desktop
For Local STDIO Mode (default):
{
"mcpServers": {
"heuris-biomcp": {
"command": "biomcp",
"env": {
"NCBI_API_KEY": "your_ncbi_api_key_here"
}
}
}
}For Remote HTTP Mode (using live demo or your own deployed server):
Use Claude's Customize > Connectors flow and enter:
https://heuris-biomcp.onrender.com/mcpIf you are connecting with an older MCP client that still expects SSE, use:
https://heuris-biomcp.onrender.com/sse💡 Tip: Get a free NCBI API key to increase rate limits from 3 to 10 requests/second.
🚀 New: Get free NVIDIA API keys for AI tools at build.nvidia.com/mit/boltz2 and build.nvidia.com/arc/evo2-40b.
Restart Claude Desktop and test:
"Search PubMed for recent papers on CAR-T cell therapy in B-cell lymphoma"
"Get the AlphaFold structure for TP53 and tell me about the confidence scores"
"What drugs are approved that target EGFR?"
"Generate a multi-omics report for KRAS"
"Predict the structure of EGFR with ligand CC1=CC=CC=C1 and compute binding affinity"
"Generate a DNA sequence starting with ATGGCG..."Usage Examples
Literature Mining
"Search PubMed for BRCA1 CRISPR correction methods published in the last 2 years"
"Find review articles about PD-1/PD-L1 immune checkpoint inhibitors"Protein Analysis
"Get UniProt info for human TP53 (P04637) including its domains and disease associations"
"Search for AlphaFold structures for insulin receptor"
"Find all PDB crystal structures of BRAF kinase domain resolved below 2.5 Ångström"Drug Discovery
"What are the top ChEMBL compounds targeting KRAS G12C mutation?"
"Get compound info for imatinib (CHEMBL941)"
"Show me gene-disease associations for BRCA1 with evidence scores"AI-Powered Structure Prediction
"Predict the 3D structure of insulin (sequence: ...) with ligand CCO"
"Compute binding affinity between EGFR and gefitinib (SMILES: ...)"
"Get structure prediction for a protein-protein complex"AI-Powered DNA Generation
"Generate a 200bp promoter sequence starting with ATG"
"Compare wildtype vs variant DNA sequence for TP53 mutation"
"Generate regulatory element for gene expression"Multi-Omics Report (Flagship)
"Generate a complete multi-omics report for EGFR"This single command queries 7 databases in parallel and returns:
Genomic location and gene summary (NCBI Gene)
Recent publications (PubMed)
Protein function and structure (UniProt + AlphaFold)
Biological pathways (Reactome)
Drug targets and clinical compounds (ChEMBL)
Disease associations with scores (Open Targets)
Expression datasets (GEO)
Active clinical trials (ClinicalTrials.gov)
v2 Extended Databases
"Get OMIM diseases associated with TP53"
"Show STRING protein interactions for EGFR"
"Get GTEx expression data for BRCA1 across tissues"
"Find mutations in TP53 from cBioPortal"
"Search GWAS for diabetes-associated SNPs"v2 Verification
"Verify the claim that TP53 is a tumor suppressor gene"
"Detect conflicts between OMIM and DisGeNET for BRCA1"v2 Experimental Design
"Generate an experimental protocol for CRISPR knockout of BRCA1"
"What cell lines should I use to study KRAS mutations?"
"Calculate sample size for detecting 2-fold change with p<0.05"v2 Session Intelligence
"What's the knowledge graph from our conversation so far?"
"Find biological connections between TP53 and EGFR"
"Export our research session as a reproducible script"
"Plan and execute a research workflow for PD-1 drug targets"Architecture
biomcp/
├── src/biomcp/
│ ├── server.py # MCP server — tool registry & dispatcher
│ ├── tools/
│ │ ├── ncbi.py # PubMed, Gene, BLAST
│ │ ├── proteins.py # UniProt, AlphaFold, PDB
│ │ ├── pathways.py # KEGG, Reactome, ChEMBL, Open Targets
│ │ ├── advanced.py # ClinicalTrials, GEO, scRNA, Ensembl,
│ │ │ # Multi-Omics, Neuroimaging, Hypothesis
│ │ ├── nvidia_nim.py # Boltz-2, Evo2-40B AI models
│ │ ├── databases.py # v2: OMIM, STRING, GTEx, cBioPortal,
│ │ │ # GWAS, DisGeNET, PharmGKB
│ │ ├── verify.py # v2: Claim verification, conflict detection
│ │ └── protocol_generator.py # v2: Experimental design tools
│ ├── core/
│ │ ├── entity_resolver.py # v2: Cross-database entity resolution
│ │ ├── knowledge_graph.py # v2: Session knowledge graph
│ │ └── query_planner.py # v2: Adaptive query planner
│ └── utils/
│ └── __init__.py # Rate limiter, cache, validators, HTTP client
├── tests/
├── pyproject.toml
└── README.mdKey design decisions:
Async-first: All API calls are fully async with
httpx, never blockingRate limiting: Token-bucket limiter per service respects each API's limits
Smart caching: TTL-based per-namespace cache (1h literature, 7d structures)
Retry logic: Exponential backoff via
tenacityfor transient failuresValidation: Input validation before any network call — never wastes API quota
Environment Variables
Variable | Description | Default |
| NCBI API key (increases rate limit to 10/s) | None (3/s) |
| NVIDIA API key for Boltz-2 structure prediction | None |
| NVIDIA API key for Evo2-40B DNA generation | None |
| Shared fallback key for both NVIDIA model integrations | None |
| BioGRID key for curated interaction queries | None |
| Transport mode: |
|
| HTTP port for hosted SSE deployments |
|
| Durable directory for saved sessions |
|
| Require auth for hosted MCP requests | disabled |
| Enable OAuth 2.1 PKCE endpoints when auth is enabled |
|
| Comma-separated API keys in | None |
| Per-key request budget per window |
|
| Per-key rate-limit window in seconds |
|
| Auto-approve the consent screen for single-user deployments |
|
| Subject bound to approved OAuth tokens |
|
| Comma-separated browser origins allowed for CORS | disabled |
| Enable per-client HTTP rate limiting |
|
| Requests allowed per rate-limit window |
|
| Rate-limit window length in seconds |
|
| Requests allowed per authenticated window |
|
| Authenticated rate-limit window length in seconds |
|
| Enable background cache warming in HTTP mode |
|
| Log level: DEBUG/INFO/WARNING/ERROR | INFO |
Get free API keys:
NVIDIA Boltz-2: https://build.nvidia.com/mit/boltz2
NVIDIA Evo2-40B: https://build.nvidia.com/arc/evo2-40b
BioGRID: https://webservice.thebiogrid.org/
Operational Endpoints
When BioMCP runs in hosted HTTP mode, these operational routes are available:
Endpoint | Purpose |
| Runtime status, HTTP policy, and session-storage configuration |
| Prometheus-compatible request, cache, latency, auth, and upstream metrics |
| Liveness and deployment metadata |
| Readiness check for orchestrators and load balancers |
| Capability-level status, including missing optional API keys |
| OAuth 2.1 metadata for hosted connectors |
| Dynamic OAuth client registration |
| OAuth authorization and PKCE consent |
| Authorization-code and refresh-token exchange |
| Streamable HTTP MCP endpoint |
| MCP SSE endpoint |
| MCP message transport endpoint |
These endpoints are designed for deployment review, Render health checks, and production smoke tests.
Contributing
Contributions are welcome! Whether it's adding a new database, fixing a bug, improving documentation, or integrating new AI models.
# Development setup
pip install -e ".[dev]"
# Run tests
pytest tests/ -v --cov=biomcp
# Lint + type check
ruff check src/
mypy src/Ideas for contributions
Add more pathway databases (Wikipathways, PathCards)
Integrate COSMIC for somatic mutations
Add protein complex data (CORUM)
Implement batch query support for high-throughput analysis
Add Jupyter notebook examples
Improve conflict resolution algorithms
Citation
If you use Heuris-BioMCP in your research, please cite:
@software{biomcp2025,
title = {Heuris-BioMCP v2: A Comprehensive MCP Server for Bioinformatics, AI Models, and Life Sciences},
year = {2025},
url = {https://github.com/SachinGawande2003/Heuris-BioMCP},
license = {MIT}
}License
MIT License — free for academic and commercial use.
Available Tools
32 toolsbiomarker_panel_designBiomarker Panel DesignARead-onlyIdempotent
Draft a disease-focused biomarker panel using Open Targets evidence with a literature fallback.
| Name | Required | Description | Default |
|---|---|---|---|
| disease | Yes | Disease, indication, or phenotype of interest. | |
| panel_size | No | Number of biomarkers to include. Default 10. | |
| context | No | Context such as oncology, inflammation, or rare disease. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, destructiveHint, idempotentHint, and openWorldHint. The description adds valuable context: the tool uses Open Targets evidence with a literature fallback, which helps agents understand behavior when evidence is lacking. This goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the core purpose and method with no unnecessary words. It is front-loaded and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters with full schema descriptions, comprehensive annotations, and no output schema, the description is fairly complete. It explains the evidence sources and fallback, though it could mention the output format. Overall, it provides sufficient context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not elaborate on parameters beyond what the schema provides, so no additional value is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('draft') and resource ('biomarker panel') and specifies the evidence sources ('Open Targets evidence with a literature fallback'). It clearly distinguishes this tool from siblings, which focus on other analyses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for disease-focused panel design but does not provide explicit when-to-use or when-not-to-use guidance, nor does it mention alternatives. It is adequate but lacks explicit direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bulk_gene_analysisBulk Gene AnalysisARead-onlyIdempotent
Analyze one or two gene panels in parallel and return a cross-gene comparison matrix. In differential mode it ranks pathways and diseases enriched in panel A versus panel B using panel-level hit fractions and fold-change-style scoring.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbols | Yes | Primary list of 2-10 HGNC gene symbols to analyze in parallel. | |
| comparison_axes | No | Aspects to compare: 'drugs', 'diseases', 'pathways', 'expression'. Default: all four. | |
| reference_gene_symbols | No | Optional reference panel of 2-10 HGNC gene symbols for differential mode. | |
| group_a_label | No | Display label for the primary panel in differential results. | |
| group_b_label | No | Display label for the reference panel in differential results. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, destructiveHint=false, idempotentHint=true, openWorldHint=true. The description adds behavioral context about the analysis process, scoring methodology, and return matrix format, which goes beyond annotations without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loads the core action, and provides essential detail on differential mode in the second sentence. No superfluous text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given five parameters, no output schema, and complexity of panel analysis, the description covers the main purpose and differential mode. However, it lacks details on output format specification (e.g., matrix structure) and constraints like gene list limits mentioned in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed parameter descriptions. The description adds high-level context about differential mode and scoring but does not provide additional semantic detail for individual parameters beyond what is in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: analyzing one or two gene panels in parallel and returning a cross-gene comparison matrix. It also distinguishes differential mode with pathway and disease enrichment ranking, which differentiates it from sibling tools like pathway_analysis or get_gene_disease_associations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains two usage modes (single panel vs differential) and when to use each. It does not explicitly state when not to use it or name alternatives, but the context from sibling tools implies its niche for bulk comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crispr_analysisCRISPR AnalysisCRead-onlyIdempotent
Merged CRISPR workflow covering guide design, guide scoring, off-target review, base editing, and repair outcome estimation.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | CRISPR workflow step. | design |
| gene_symbol | No | Target gene symbol. | |
| guide_sequence | No | Guide sequence for scoring, off-target review, or repair analysis. | |
| target_mutation | No | Desired mutation for base-edit design. | |
| repair_template | No | Optional HDR repair template. | |
| cas_variant | No | CRISPR nuclease. | SpCas9 |
| target_region | No | Guide search region, such as 'early_exons' or 'all_coding'. | |
| n_guides | No | Number of guides to return. Default 5. | |
| min_score | No | Minimum guide score for design. | |
| mismatches | No | Maximum mismatches for off-target review. Default 3. | |
| genome | No | Genome assembly used for off-target review. | hg38 |
| cell_line | No | Cell-line context for repair estimates. | |
| use_blast | No | Use BLAST to supplement off-target review. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is clear. However, the description adds no additional behavioral context, such as whether results are cached, computational cost, or prerequisites (e.g., genome availability). It merely states the covered steps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that encapsulates the tool's scope without redundant wording. It is front-loaded with the key term 'CRISPR workflow' and then enumerates steps, making it efficient. Slightly longer due to listing all steps, but still concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (13 parameters, 1 required) and the absence of an output schema, the description is minimally adequate. It explains the high-level purpose but does not guide which parameters correspond to which step or provide examples beyond the schema example. The schema's examples partly compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3. The description does not add meaning beyond the schema, but the schema itself is thorough, including enums and defaults. The parameter 'action' is implicitly linked to the workflow steps, but this is already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly lists the CRISPR workflow steps (design, scoring, off-target, base editing, repair), making the tool's purpose explicit. However, it does not differentiate from sibling tools like 'run_blast' or 'variant_analysis', which are distinct but could be confused without context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives (e.g., 'run_blast' for sequence alignment). The description lacks context for tool selection or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drug_interaction_checkerDrug Interaction CheckerBRead-onlyIdempotent
Check FDA label interaction context between a primary drug and a list of co-medications.
| Name | Required | Description | Default |
|---|---|---|---|
| drug_name | Yes | Primary drug name. | |
| co_medications | No | Co-medications to screen against the primary label. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds the specific data source (FDA labels) and the type of check, but does not mention limitations or output behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single, clear sentence without redundancy. Front-loaded and appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with two parameters and no output schema, the description is mostly adequate but lacks explanation of what 'interaction context' means in the output (e.g., severity, management). A minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are described. The description repeats 'primary drug' and 'co-medications' without adding format constraints or conventions. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks FDA label interaction context between a primary drug and co-medications. The verb 'Check' and resource 'FDA label interaction context' are specific. It distinguishes from siblings like drug_safety and pharmacogenomics_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like drug_safety. The description implies use for interaction context but lacks 'when to use' or 'when not to use' statements.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drug_safetyDrug SafetyARead-onlyIdempotent
Merged FDA drug-safety workflow for adverse-event search, signal detection, label review, and head-to-head comparison.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Drug-safety workflow step. | events |
| drug_name | Yes | Generic or brand drug name. | |
| comparator_drug | No | Comparator drug used when action='compare'. | |
| event_type | No | Safety category filter. | all |
| serious_only | No | Restrict adverse-event search to serious reports. | |
| event_terms | No | Optional adverse-event terms for signal detection. | |
| max_results | No | Maximum reports to summarize. Default 50. | |
| patient_sex | No | Patient sex filter. | |
| age_group | No | Age filter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true. The description adds the 'merged' nature, suggesting aggregation from multiple FDA sources. However, it does not disclose other behavioral traits such as rate limits, data freshness, or how the merged workflow is coordinated. The description adds marginal value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the main purpose and enumerates the core actions. Every phrase contributes to understanding, with no fluff. It is optimally concise for a multi-action tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the schema covering parameters 100%, the description fails to explain how the various actions work, what the output looks like, or how to combine them in a workflow. The openWorldHint suggests external data, but no output schema or description of return format is provided. For a tool with 9 parameters and multiple modes, the description is too short to give a complete picture, leaving significant gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add any parameter-specific information beyond what the schema already provides. It merely lists the high-level workflow steps, which are mapped to the 'action' parameter's enum. No additional semantic context for parameters like 'serious_only' or 'age_group' is given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a merged FDA drug-safety workflow covering four specific tasks: adverse-event search, signal detection, label review, and head-to-head comparison. This directly distinguishes it from sibling tools like drug_interaction_checker and the many gene/sequence tools, as it uniquely addresses multi-step drug safety analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for combined drug-safety tasks but does not explicitly state when to use it versus alternatives like drug_interaction_checker. The context signals (sibling list) and the name 'drug_safety' provide enough cues for an agent to select it for safety workflows, but a brief note on exclusion would improve clarity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_proteinFind ProteinARead-onlyIdempotent
Unified protein discovery across UniProt and PDB. Use accession for a direct record lookup or query to search reviewed proteins and experimental structures.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Protein, gene, family, or free-text search query. | |
| source | No | Which source to search. | auto |
| accession | No | Optional UniProt accession for direct lookup. | |
| organism | No | Organism filter for UniProt search. Default 'homo sapiens'. | |
| reviewed_only | No | Limit UniProt search to reviewed Swiss-Prot entries. | |
| max_results | No | Maximum results per source. Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true. The description adds minimal behavioral context beyond stating the unified search scope. It does not contradict annotations, but fails to provide additional traits like rate limits or response structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded: two sentences packing the core purpose and usage modes without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters, no output schema, and good annotations, the description covers the essential use cases and hints at the two search strategies. It lacks details on how parameters like 'source' or 'organism' interact, but these are documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is well-documented. The description adds high-level context about the two main use modes (accession vs query), but does not elaborate on specifics like the 'source' or 'organism' parameters beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Unified protein discovery across UniProt and PDB'. It specifies the two main use modes (accession lookup or query search), and distinguishes it from sibling tools like get_protein_info or get_alphafold_structure which target specific data retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear instructions on when to use each mode: 'Use accession for a direct record lookup or query to search reviewed proteins and experimental structures.' It implicitly guides the agent on choosing between the two input types, though it does not explicitly exclude scenarios or mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_repurposing_candidatesFind Repurposing CandidatesARead-onlyIdempotent
Drug repurposing engine: surface approved drugs with activity against a target/disease.
| Name | Required | Description | Default |
|---|---|---|---|
| disease | Yes | Target disease. | |
| gene_target | No | Primary gene target. Optional. | |
| max_candidates | No | Maximum repurposing candidates Default 15. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and other traits. The description adds minimal behavioral context beyond 'surface approved drugs,' but does not contradict annotations. More detail on algorithm or data source would improve transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long with no extraneous information. It is front-loaded and efficiently conveys the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite good annotations, the description lacks details on output format, scoring or prioritization of candidates, and limitations. For a discovery tool with no output schema, more context is needed to fully inform an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description does not add new meaning beyond the schema, resulting in a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'surface' and resource 'approved drugs' with a clear context of activity against a target/disease. It effectively distinguishes from sibling tools like drug_interaction_checker and get_drug_targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for drug repurposing but does not provide explicit guidance on when to use this tool versus alternatives. No exclusions or when-not scenarios are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_dna_evo2Generate DNA Evo2ARead-only
Evo2 DNA workflow. Use mode='generate' for sequence generation or mode='score' for wildtype-versus-variant scoring.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Evo2 workflow mode. | generate |
| sequence | No | Seed DNA sequence for generation. | |
| num_tokens | No | Number of DNA tokens to generate. Default 200. | |
| temperature | No | Sampling temperature. | |
| top_k | No | Top-K sampling parameter. Default 4. | |
| top_p | No | Top-P sampling parameter. | |
| enable_logits | No | Return per-token logits for generation. | |
| num_generations | No | Independent generation runs. Default 1. | |
| wildtype_sequence | No | Reference DNA sequence used when mode='score'. | |
| variant_sequence | No | Variant DNA sequence used when mode='score'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds the workflow nature (generation vs scoring) but does not disclose any additional behavioral traits such as resource usage, API costs, or limitations. The bar is lowered due to annotations, but still basic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with a single sentence that covers the core functionality. It could be front-loaded with an action verb but is efficient for the main purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having 10 parameters and no output schema, the description is extremely brief. It does not explain what the tool returns, how to interpret results, or any necessary context for operation. The agent lacks guidance on expected output format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents all 10 parameters. The description adds value by explaining the functional difference between the two modes (generate vs score), but does not elaborate on other parameters like num_tokens or temperature, which are self-explanatory in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it is an Evo2 DNA workflow with two distinct modes: generate and score. It separates generation from scoring, which distinguishes it from sibling tools that are focused on analysis, not DNA sequence generation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit instructions on when to use each mode via the 'mode' parameter. However, it does not mention any prerequisites, exclusions, or alternatives to guide the agent in choosing this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
genome_browser_snapshotGenome Browser SnapshotARead-onlyIdempotent
Generate genome-browser links and locus context for a gene or explicit genomic interval.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbol | No | Gene symbol to resolve to a locus. | |
| region | No | Explicit region such as 'chr17:43044295-43170245'. | |
| flank_bp | No | Flanking sequence to include around the locus. Default 25000. | |
| assembly | No | Genome assembly. | GRCh38 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, covering safety. The description adds value by specifying the generated output (links and context), providing behavioral context beyond what annotations offer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with no extraneous words. It efficiently communicates the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity, clear description, complete schema, and detailed annotations, the description is mostly complete. However, the term 'locus context' could be interpreted as ambiguous, and there is no mention of return format or that it generates URLs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% description coverage so baseline is 3. The description adds context that the tool works for a gene or genomic interval, tying the parameters together, which is helpful beyond the per-parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'generate' and the resources 'genome-browser links and locus context', and it distinguishes from sibling tools by specifying the output type (links and context) rather than analysis or other operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is provided. The usage context is implied by the description but not compared to alternatives, which limits the agent's ability to decide when to use this tool over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_alphafold_structureGet AlphaFold StructureARead-onlyIdempotent
AlphaFold DB predicted structure: per-residue pLDDT confidence stats, PDB/mmCIF download URLs. pLDDT ≥90=very high, 70–90=confident, <50=disordered.
| Name | Required | Description | Default |
|---|---|---|---|
| uniprot_accession | Yes | UniProt accession (e.g. 'P04637'). | |
| model_version | No | AlphaFold model version. Default: v4. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, indicating a safe, idempotent read operation. The description adds value by detailing the specific outputs (per-residue stats, download URLs) and providing confidence thresholds, which go beyond the annotations. No contradictions found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loading the core purpose and immediately adding interpretive context (confidence thresholds). Every sentence provides necessary information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description explains what is returned (stats and URLs) and how to interpret pLDDT scores. It covers the main behavioral aspects given the rich annotations (readOnly, idempotent). However, it does not mention error cases (e.g., missing accession) or response format, but overall it is sufficiently complete for the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions for both parameters (uniprot_accession and model_version). The description does not elaborate on parameter syntax or behavior but implies the output depends on the accession. Baseline 3 is appropriate since the schema carries the burden adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves AlphaFold predicted structure data, including per-residue pLDDT confidence statistics and download URLs for PDB/mmCIF formats. This specific verb-resource pair distinguishes it from sibling tools like predict_structure_boltz2, which generates structures rather than retrieving existing ones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides confidence thresholds (pLDDT ≥90, 70–90, <50) to aid interpretation but does not explicitly state when to use this tool versus alternatives like predict_structure_boltz2 (for predictions) or other retrieval tools. No when-not or exclusion criteria are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_drug_targetsGet Drug TargetsARead-onlyIdempotent
ChEMBL drug-target activities: IC50, Ki, Kd values, assay types, approval status. Auto-indexes drug→gene edges into knowledge graph.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbol | Yes | Target gene symbol (e.g. 'EGFR', 'BRAF', 'KRAS'). | |
| max_results | No | Drug entries Default 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, openWorld), description adds that the tool auto-indexes edges into a knowledge graph and specifies the types of values returned, providing useful behavioral context. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: first lists key outputs, second mentions an important side-effect (auto-indexing). Every part adds value, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Description is adequate for a simple query tool with good annotations, but missing explicit details about output structure (e.g., list of objects with fields like IC50, Ki). Given no output schema, more detail would help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already describes both parameters well (100% coverage). Description adds value by clarifying that the output includes specific activity types and metadata, helping interpret the results expected from the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it retrieves ChEMBL drug-target activities (IC50, Ki, Kd, assay types, approval status) for a given gene, distinguishing it from siblings that focus on other aspects like drug interactions or disease associations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for obtaining drug-target activity data but does not explicitly state when to use this tool versus alternatives like drug_interaction_checker, leaving the agent to infer context from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_gene_disease_associationsGet Gene Disease AssociationsARead-onlyIdempotent
Open Targets gene-disease evidence across 6 datatypes: genetic_association, somatic_mutation, known_drug, animal_model, affected_pathway, literature. Auto-indexes gene→disease edges into knowledge graph.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbol | Yes | HGNC gene symbol. | |
| max_results | No | Associations Default 15. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint, idempotentHint, etc.) already indicate safety and idempotency. The description adds context about auto-indexing into a knowledge graph, which is beyond what annotations provide. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences efficiently convey purpose and a behavioral detail. The second sentence about auto-indexing is somewhat extraneous but not harmful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers data sources but omits return value structure or pagination behavior. Without an output schema, the agent lacks details on what fields are returned, which is a notable gap for a retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% description coverage for both parameters (gene_symbol, max_results) with clear details like HGNC and defaults. Description does not add further parameter semantics beyond listing datatypes, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it retrieves Open Targets gene-disease evidence across 6 specified datatypes, making it distinct from sibling tools like get_gene_info or search_gwas_catalog.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes the tool's function but does not explicitly state when to use it versus alternatives such as get_gene_info or variant_analysis. Usage context is implied but not clearly delineated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_gene_infoGet Gene InfoARead-onlyIdempotent
Retrieve gene information from NCBI Gene — symbol, full name, chromosomal location, aliases, RefSeq IDs, and functional summary. Auto-indexes gene entity into session knowledge graph.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbol | Yes | HGNC gene symbol (e.g. 'TP53', 'BRCA1', 'EGFR'). | |
| organism | No | Species. Default: 'homo sapiens'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, destructiveHint, idempotentHint, openWorldHint. The description adds valuable behavioral context: it auto-indexes the gene entity into the session knowledge graph (a side effect beyond read-only). No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with purpose, no wasted words. Every sentence provides value (purpose and data types; auto-indexing side effect).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description lists return data types (symbol, full name, etc.). Annotations cover behavioral aspects. Complexity is moderate; the description is sufficiently complete for a retrieval tool, though explicit return format could help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with good parameter descriptions. The description adds that gene_symbol is an HGNC symbol and specifies the default organism, but these are already in the schema. It does not significantly enhance understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a specific verb ('Retrieve') and resource ('gene information from NCBI Gene') and lists concrete data types. It distinguishes from siblings like 'get_protein_info' by focusing on NCBI Gene and auto-indexing into the session knowledge graph.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for gene information retrieval but does not explicitly state when to use this tool vs. alternatives (e.g., 'find_protein' or 'get_protein_info'). No exclusions or comparison provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_protein_infoGet Protein InfoARead-onlyIdempotent
Full UniProt Swiss-Prot entry: function, domains, PTMs, GO terms, disease links, sequence. Prefer reviewed accessions (P/Q/O prefix). Auto-indexes protein + disease edges into session knowledge graph.
| Name | Required | Description | Default |
|---|---|---|---|
| accession | Yes | UniProt accession (e.g. 'P04637' for human TP53). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description claims the tool auto-indexes into the session knowledge graph, contradicting the readOnlyHint annotation, which indicates no side effects. This is an annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads key content types, and includes a usage hint. Every word is informative with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lists the types of data returned (function, domains, etc.), compensating for the lack of an output schema. However, the contradiction reduces confidence in behavioral completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema fully covers the single parameter with an example. The description adds value by recommending reviewed accessions (P/Q/O prefix), providing guidance beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves a full UniProt Swiss-Prot entry with specific content types (function, domains, PTMs, GO terms, disease links, sequence), distinguishing it from sibling tools like get_gene_info and find_protein.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises preferring reviewed accessions (P/Q/O prefix) but does not specify when to use this tool versus alternatives like find_protein or get_gene_info. Usage context is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
multi_omics_gene_reportMulti Omics Gene ReportARead-only
FLAGSHIP: 7-database parallel integration — NCBI Gene, PubMed, Reactome, ChEMBL, Open Targets, GEO, ClinicalTrials.gov. One call, complete overview.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbol | Yes | HGNC gene symbol (e.g. 'EGFR', 'TP53', 'BRCA1', 'KRAS'). | |
| detail_level | No | Response size budget. Use 'compact' for connector-friendly summaries, 'standard' for the default layer payloads, and 'full' to preserve full layer outputs. | compact |
| include_synthesis | No | When true, add a Claude-written synthesis if ANTHROPIC_API_KEY is configured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and destructiveHint=false. The description adds value by listing the specific databases integrated and stating the tool produces a 'complete overview', which augments the behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The 'FLAGSHIP' label front-loads importance, and every word serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 databases, 3 parameters, no output schema), the description adequately explains scope and intent. Though it does not detail the report's structure, the integration of databases is sufficient for agent selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already well-documented. The description does not add additional parameter-level meaning beyond what the schema provides, meeting the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as a 'FLAGSHIP' that integrates 7 databases (NCBIGene, PubMed, etc.) for a complete gene overview. It distinguishes itself from sibling tools by emphasizing parallel multi-database integration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a comprehensive multi-database gene report is needed via the 'FLAGSHIP' label and database list, but does not explicitly state when to use or avoid this tool versus alternatives like get_gene_info or search_pubmed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
network_enrichmentNetwork EnrichmentARead-onlyIdempotent
Summarize recurrent Reactome pathways and STRING network hubs across a gene set.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_list | Yes | Input genes for enrichment analysis. | |
| min_string_score | No | Minimum STRING confidence score. Default 700. | |
| max_results | No | Maximum pathways and hubs to return. Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, which adequately inform about safety and idempotency. The description adds no additional behavioral insights (e.g., rate limits, data freshness). It repeats the purpose but does not expand beyond annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise and front-loaded, clearly communicating the tool's purpose without extraneous words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a simple interface with 3 parameters and no output schema. The description covers the core function but omits what the output looks like (e.g., list of pathways with scores, network nodes). The input example partially compensates. Overall, it is adequate but could mention return format for completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% as all three parameters have descriptions in the input schema. The tool description does not add any extra meaning or nuances beyond what is already in the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'summarize' and identifies distinct resources (Reactome pathways, STRING network hubs) applied to a gene set. It clearly distinguishes from sibling tools like 'pathway_analysis' which likely performs standard enrichment, and 'find_protein' which focuses on individual proteins.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for summarizing network-level insights, but it does not explicitly state when to use it versus alternatives such as 'pathway_analysis' or 'verify_biological_claim'. No when-not or clear context for selection is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pathway_analysisPathway AnalysisBRead-onlyIdempotent
Merged pathway workflow. Search KEGG, retrieve genes for a pathway, or assemble Reactome plus KEGG context for a gene.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | Workflow step. | auto |
| db | No | Preferred pathway database. | auto |
| query | No | Free-text pathway search term. | |
| gene_symbol | No | Gene symbol for Reactome or KEGG context. | |
| pathway_id | No | KEGG pathway identifier such as 'hsa05200'. | |
| organism | No | KEGG organism code. Default 'hsa'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering safety and idempotency. The description adds no further behavioral context (e.g., data source freshness, failure modes, or that it combines two databases). Given the strong annotation coverage, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that communicates core capabilities without fluff. However, it lacks front-loading of the most critical information (the action parameter) and could benefit from a more structured format (e.g., listing actions).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and 6 parameters with multiple actions, yet the description does not explain what the tool returns or describe the 'auto' action. This leaves the agent without crucial context about invocation results and action selection, making the description incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all parameters have descriptions in the schema. The tool description does not add additional meaning beyond what the schema provides (e.g., it mentions actions but does not clarify which action to use for which scenario). Per guidelines, baseline is 3 with high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists three specific capabilities (search KEGG, retrieve genes for a pathway, assemble Reactome plus KEGG context) tied to actions in the schema. It clearly identifies the tool as a merged pathway workflow, but the term 'merged pathway workflow' is somewhat vague and the 'auto' action is not described, slightly reducing clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With 30+ sibling tools including 'network_enrichment' and 'get_gene_info', explicit when-to-use or when-not-to-use guidance is missing, limiting the agent's ability to select the correct tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pharmacogenomics_reportPharmacogenomics ReportBRead-onlyIdempotent
Summarize CPIC-style pharmacogenomic genes and supporting PGx evidence for a drug.
| Name | Required | Description | Default |
|---|---|---|---|
| drug_name | Yes | Drug of interest. | |
| gene_symbol | No | Optional gene to force into the report. | |
| max_annotations | No | Maximum supporting annotations per gene. Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, so the description does not need to repeat these. It adds the context 'CPIC-style' and 'PGx evidence', which is useful but minimal. No additional behavioral traits (e.g., rate limits, auth) are disclosed, but the annotation coverage is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that is front-loaded with the core action. It is concise without being overly terse, but there is room to add a bit more context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters, no output schema, and good annotations, the description is adequate but not rich. It does not explain what the output looks like or what 'CPIC-style' entails exactly. More detail on the report structure would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all parameters. The description does not add meaning beyond what is in the schema (e.g., 'Drug of interest', 'Optional gene to force into the report'). Baseline 3 is appropriate as the description offers no extra clarification.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Summarize CPIC-style pharmacogenomic genes and supporting PGx evidence for a drug.' It uses a specific verb ('Summarize') and resource ('pharmacogenomic genes and evidence'), and the mention of 'CPIC-style' distinguishes it from general drug info tools like drug_interaction_checker or drug_safety.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is provided. The description does not mention alternatives or context for selecting this tool over siblings such as drug_safety, get_drug_targets, or drug_interaction_checker. The agent must infer usage from the tool name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_structure_boltz2Predict Structure Boltz2BRead-only
Boltz-2 structure workflow. Use mode='structure' for direct multimolecular structure prediction or mode='protein_ligand' for the integrated UniProt-to-docking workflow.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Boltz-2 workflow mode. | structure |
| protein_sequences | No | Protein sequences for direct Boltz-2 prediction. | |
| ligand_smiles | No | Ligand SMILES strings. | |
| dna_sequences | No | Optional DNA sequences for complex prediction. | |
| rna_sequences | No | Optional RNA sequences for complex prediction. | |
| uniprot_accession | No | UniProt accession used when mode='protein_ligand'. | |
| predict_affinity | No | Predict ligand binding affinity when ligands are present. | |
| method_conditioning | No | Optional structure-style conditioning. | x-ray |
| pocket_residues | No | Optional binding-pocket residue constraints. | |
| recycling_steps | No | Boltz-2 recycling iterations. Default 3. | |
| sampling_steps | No | Diffusion sampling steps. Default 200. | |
| diffusion_samples | No | Number of structural hypotheses. Default 1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and description adds mode context. However, it does not disclose limitations, input conflicts, or output behavior beyond schema info.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two clear sentences that are front-loaded and efficient, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters and no output schema, the description is too sparse. It does not cover parameter semantics for complex inputs like pocket_residues or the interaction between multiple sequence types.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline 3. The description adds minimal meaning beyond schema by explaining modes, but does not clarify parameter interactions (e.g., protein_sequences vs uniprot_accession).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a Boltz-2 structure workflow with two modes (structure and protein_ligand), distinguishing it from sibling tools like get_alphafold_structure. However, it could be more specific about the output format.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description instructs which mode to use for multimolecular vs. protein-ligand docking, but lacks explicit guidance on when not to use this tool or comparisons to alternatives like AlphaFold.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
protein_binding_pocketProtein Binding PocketARead-onlyIdempotent
Summarize candidate binding sites from UniProt functional-site annotations plus AlphaFold confidence context.
| Name | Required | Description | Default |
|---|---|---|---|
| accession | No | UniProt accession. | |
| query | No | Protein or gene query used to resolve a reviewed accession. | |
| max_sites | No | Maximum candidate sites to return. Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true. The description adds valuable context about data sources (UniProt and AlphaFold) and the summarization behavior, enhancing understanding beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that efficiently conveys the tool's purpose without unnecessary words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters (all optional), no output schema, and annotations indicating read-only behavior, the description adequately covers what the tool does and its data sources. However, it lacks details on output format or what 'summarize' entails.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and all parameters are well-described in the schema. The description adds general context about the tool's output but does not provide specific additional meaning for the parameters beyond what the schema includes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'summarize', the resource 'candidate binding sites', and the data sources 'UniProt functional-site annotations plus AlphaFold confidence context'. It effectively distinguishes this tool from siblings like get_alphafold_structure or find_protein.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for summarizing binding sites but does not explicitly state when to use this tool vs. alternatives, nor does it provide when-not-to-use guidance or mention prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
protein_family_analysisProtein Family AnalysisARead-onlyIdempotent
Summarize protein family and domain context from curated UniProt annotations with direct Pfam and InterPro links.
| Name | Required | Description | Default |
|---|---|---|---|
| accession | No | UniProt accession. | |
| query | No | Protein or gene query used to resolve an accession. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, etc. The description adds context by specifying that data comes from 'curated UniProt annotations' and that output includes 'direct Pfam and InterPro links', which informs the agent about the source and format of the results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no redundant information. Every word contributes to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description provides a general sense of what the tool returns (summaries with links) but lacks details on the structure, pagination, or how the two optional parameters are used together. The example helps but does not fully clarify behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both parameters have descriptions). The description does not add additional meaning beyond the schema; it restates the general purpose without clarifying parameter-specific semantics like how accession and query interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Summarize', the resource 'protein family and domain context from curated UniProt annotations', and the added value 'with direct Pfam and InterPro links'. It distinguishes from sibling tools like get_protein_info or pathway_analysis by specifying the focus on family/domain context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for obtaining protein family and domain summaries but does not explicitly state when to use this tool versus alternatives like get_protein_info or network_enrichment. No exclusions or when-not conditions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rare_disease_diagnosisRare Disease DiagnosisARead-onlyIdempotent
Normalize phenotype terms and rank OMIM differentials for a candidate gene.
| Name | Required | Description | Default |
|---|---|---|---|
| phenotype_terms | No | Phenotype terms or symptoms to normalize. | |
| gene_symbol | No | Candidate gene symbol. | |
| max_results | No | Maximum ranked differentials to return. Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, etc. The description adds that the tool normalizes terms and ranks differentials, but doesn't disclose additional behavioral traits beyond what annotations provide. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, efficient sentence that is front-loaded and contains no wasted words. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description adequately explains the tool's purpose but does not describe the output format or ranking criteria. With no output schema, additional detail on return structure would improve completeness for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with descriptions for each parameter. The description does not add new parameter-level information beyond schema, so baseline 3 is appropriate. The mention of 'normalize' adds some context but not parameter-specific meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool normalizes phenotype terms and ranks OMIM differentials for a candidate gene, using specific verbs and resources. It distinguishes from sibling tools like get_gene_disease_associations which likely list all associations without ranking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for rare disease diagnosis with a candidate gene, but lacks explicit guidance on when not to use it or mention of alternatives. However, the context of siblings helps infer its specific role.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rnaseq_deconvolutionRNA-seq DeconvolutionARead-onlyIdempotent
Marker-based heuristic deconvolution of a bulk RNA-seq profile into likely cell-type fractions.
| Name | Required | Description | Default |
|---|---|---|---|
| expression_profile | No | Expression profile keyed by gene symbol with numeric abundance values. | |
| ranked_genes | No | Optional ranked marker genes when numeric expression is unavailable. | |
| max_cell_types | No | Maximum cell types to return. Default 5. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readonly, non-destructive, idempotent behavior. The description adds that the method is 'heuristic' and 'likely,' informing the agent about approximate nature. This adds value beyond annotations, though more detail on limitations would be beneficial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that conveys the essential purpose without any extraneous words. It is appropriately front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description does not explain what the tool returns (e.g., cell-type fractions with confidence scores). The example in the schema helps but is not in the description. Given the tool's complexity and three parameters, the description should cover return format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all three parameters with clear descriptions, achieving 100% coverage. The tool description does not add additional meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs deconvolution of bulk RNA-seq data into cell-type fractions, using a marker-based heuristic. This verb+resource combination is specific and distinguishes it from sibling tools like bulk_gene_analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives, nor does it mention prerequisites or scenarios where it is inappropriate. The agent is left to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_blastRun BlastARead-only
Run NCBI BLAST sequence alignment (blastp/blastn/blastx/tblastn). Async polling — waits up to 120s for results.
| Name | Required | Description | Default |
|---|---|---|---|
| sequence | Yes | Amino acid or nucleotide sequence (raw or FASTA). | |
| program | No | BLAST program. | blastp |
| database | No | Target database. | nr |
| max_hits | No | Alignments to return Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only and non-destructive. The description adds the key behavioral trait of async polling with a 120s wait, which is valuable beyond annotations. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loads the core purpose, and every word adds value. No redundant or extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 4 parameters, all documented, and no output schema. The description does not describe the return format or how results are structured, which could be helpful for an agent. However, the tool is relatively straightforward and the async behavior is noted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add additional parameter meaning beyond what the schema provides. It mentions the programs but does not explain them further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs NCBI BLAST sequence alignment and lists the supported programs (blastp/blastn/blastx/tblastn). This gives a specific verb and resource, and distinguishes from sibling tools which are primarily gene/drug/structure analysis tools, not sequence alignment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions async polling with a 120s timeout, but does not provide explicit guidance on when to use this tool versus alternatives or when not to use it. The context of sibling tools implies it is the only BLAST tool, so usage context is clear but no exclusions or alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_cbio_mutationsSearch cBio MutationsBRead-onlyIdempotent
cBioPortal cancer mutation frequencies across TCGA cohorts.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbol | Yes | HGNC gene symbol. | |
| cancer_type | No | TCGA cancer type (e.g. 'luad'). Empty=pan-cancer. | |
| max_studies | No | Studies to query Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the description adds minimal safety context. It does clarify the data source (cBioPortal) and scope (TCGA), which provides some behavioral context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that concisely conveys the tool's purpose with no unnecessary words. Excellent structure for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters and no output schema, the description is too brief. It lacks information about the return format, pagination, rate limits, or how results are structured, leaving significant gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents parameters. The description adds no additional meaning beyond what's in the schema, meeting the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool queries cBioPortal for cancer mutation frequencies across TCGA cohorts. This is specific and distinguishes it from sibling tools like bulk_gene_analysis or variant_analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool vs alternatives or when not to use it. While the context implies use for mutation frequencies, it lacks any caveats or comparative advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_clinical_trialsSearch Clinical TrialsARead-onlyIdempotent
ClinicalTrials.gov v2: trial status, phase, interventions, enrollment, eligibility. Auto-indexes drug→disease treatment edges from trials.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Disease, drug, gene, or condition. | |
| status | No | Trial status. | RECRUITING |
| phase | No | Phase filter (optional). | |
| max_results | No | Results Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, openWorldHint=true, so the safety profile is clear. The description adds the behavioral trait of 'auto-indexes drug→disease treatment edges', which provides extra context beyond annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences that are concise and front-loaded. First sentence states the core functionality and filters; second sentence adds a notable feature. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 4 parameters, good annotations, and no output schema, the description covers the core purpose and a distinguishing feature. It lacks explicit return format details, but the schema examples partially compensate. Overall adequate for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters adequately. The description does not add new parameter-specific meaning but reinforces the concept of trial filters. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it searches ClinicalTrials.gov with filters for status, phase, interventions, enrollment, eligibility. It distinguishes from sibling tools (e.g., search_pubmed, get_gene_info) by focusing on clinical trials, and the explicit mention of 'drug→disease treatment edges' adds specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for finding clinical trials but does not provide explicit when-to-use or alternatives. However, the context of sibling tools (none of which are clinical trial search) makes the purpose clear, and the example in the schema provides usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_gwas_catalogSearch GWAS CatalogBRead-onlyIdempotent
NHGRI-EBI GWAS Catalog: genome-wide significant associations for a gene.
| Name | Required | Description | Default |
|---|---|---|---|
| gene_symbol | Yes | HGNC gene symbol. | |
| max_results | No | Associations Default 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering safety. The description adds minimal behavioral context beyond stating the data source and scope. It does not describe pagination, result format, or any side effects, but the annotations largely compensate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, concise and to the point. It front-loads the essential information (source and what is retrieved). No extraneous words. Could include a brief note about the optional parameter, but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with 2 parameters and no output schema, the description is mostly adequate but lacks details about the return format, data fields, or necessary credentials. Users unfamiliar with the GWAS Catalog may need more context. The example in input schema helps, but the description itself is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both gene_symbol and max_results are described in the schema. The description does not add any parameter information, but the schema already explains them adequately. Baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it retrieves genome-wide significant associations from the NHGRI-EBI GWAS Catalog for a given gene. It uses specific resource (GWAS Catalog) and outcome (associations), but lacks an explicit verb like 'search' or 'retrieve'. It differentiates from siblings by specifying the data source, but not exhaustively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like get_gene_disease_associations or find_protein. The description does not mention prerequisites, exclusions, or when not to use it. The context from sibling tool names implies a niche, but explicit guidance is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_pubmedSearch PubMedARead-onlyIdempotent
Search PubMed for scientific literature. Supports full NCBI query syntax (MeSH terms, Boolean operators, field tags, date ranges). Returns articles with title, authors, abstract, DOI, PMID, journal, year, and MeSH terms. Results auto-indexed into session knowledge graph.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | PubMed query. E.g. 'BRCA1[Gene] AND breast cancer AND Review[pt]' | |
| max_results | No | Articles to return Default 10. | |
| sort | No | Sort order. | relevance |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world. The description adds that results are auto-indexed into the session knowledge graph and lists return fields. No contradiction; adds useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences that front-load the primary purpose, then add capabilities and return info. No fluff; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with no output schema, the description covers purpose, features, return fields, and a side effect. It does not mention pagination or limits, but is otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed parameter descriptions. The description does not add additional meaning beyond what the schema provides, thus baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches PubMed for scientific literature, specifies it supports full NCBI query syntax, and lists return fields. It distinguishes from sibling tools like search_clinical_trials by being specifically for PubMed literature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for literature searches and mentions advanced query features, but does not explicitly compare to sibling tools or provide when-to-use guidance. The context is clear though not exclusive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionSessionCDestructive
Merged research-session workflow for entity resolution, live graph inspection, persisted graph save/restore via MCP resources, provenance export, and automated planning.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Session workflow step. | resolve_entity |
| query | No | Entity query or planning goal. | |
| hint_type | No | Entity type hint used for resolution. | |
| goal | No | Explicit research goal for planning. | |
| depth | No | Planning depth. | standard |
| min_path_length | No | Minimum path length for unexpected connections. Default 2. | |
| session_id | No | Saved session identifier used for restore or explicit save naming. | |
| label | No | Human-readable label for saved sessions. | |
| merge | No | Merge a restored session into the current live graph instead of replacing it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true. Description adds context about merge semantics and MCP resource usage, but does not detail specific behavioral nuances like session persistence boundaries or side effects of each action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and lists key capabilities. Could be more front-loaded with the core purpose, but no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, 12 actions, and no output schema, the description is too sparse. It does not detail each action's behavior, expected inputs, or outcomes, leaving critical gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add additional parameter-level context beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists multiple capabilities (entity resolution, graph inspection, save/restore, export, planning) but fails to define a single specific verb+resource. It distinguishes from sibling tools as a session management tool, but the purpose is broad and lacks focus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus sibling tools like find_protein or pathway_analysis. No 'when to use' or 'when not to use' rationale, leaving the agent to infer from the tool's broad functionality.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
structural_similarityStructural SimilarityARead-onlyIdempotent
PubChem-backed structural similarity search from a compound name or SMILES string.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Compound name or identifier. | |
| smiles | No | Canonical or query SMILES string. | |
| threshold | No | PubChem 2D similarity threshold. Default 90. | |
| max_results | No | Maximum similar compounds to return. Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safe read-only nature is clear. The description adds that the search is 'PubChem-backed' and uses a threshold, which provides context beyond annotations. However, it does not discuss rate limits, data freshness, or limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 11 words, front-loading the key information ('PubChem-backed structural similarity search'). Every word is meaningful and there is no wasted content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters and no output schema, the description tells what it does and the input types. It is mostly complete for a search tool with rich annotations, though it could hint at the return format (e.g., list of similar compounds with scores).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds minimal extra meaning beyond what the schema provides (e.g., mentioning 'compound name or SMILES string' maps to query and smiles, but the schema already has descriptions). Per rules, baseline is 3 when coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs a 'structural similarity search' using 'PubChem' from 'a compound name or SMILES string'. The verb 'search' and resource 'structural similarity' are specific, and it distinguishes from sibling tools which focus on other bioinformatics tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for finding similar compounds but provides no explicit guidance on when to use this tool versus alternatives (e.g., when to use a different search or analysis tool). There is no 'when not to use' or 'see also' mention.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
variant_analysisVariant AnalysisARead-onlyIdempotent
Merged variant-interpretation workflow for ACMG classification, population frequency, ClinVar review, splice review, and full integrated reporting.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Variant workflow step. | full_report |
| gene_symbol | No | HGNC gene symbol. | |
| variant | No | Variant notation, rsID, HGVS, or protein-change string. | |
| inheritance | No | Inheritance mode for ACMG scoring. | unknown |
| consequence | No | Optional VEP consequence if already known. | |
| proband_phenotype | No | Clinical phenotype context. | |
| populations | No | Population IDs to report for gnomAD frequencies. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds context about the workflow steps (ACMG, frequency, ClinVar, splice, reporting) and implies a read-only analysis. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that front-loads the key purpose ('Merged variant-interpretation workflow'). Efficient but could slightly expand on key steps without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, no output schema), the description is adequate for understanding the main purpose but does not explain return values or behavior after execution. Lacks details on output format or next steps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear parameter descriptions. The description does not add additional meaning beyond the schema, which is acceptable. Baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly identifies the tool as a merged variant-interpretation workflow for ACMG classification, population frequency, ClinVar review, splice review, and full reporting. It uses specific verbs and resources, and the scope differentiates it from sibling tools like 'bulk_gene_analysis' or 'pathway_analysis'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly suggests usage for variant interpretation but does not explicitly state when to use this tool versus alternatives like 'rare_disease_diagnosis' or 'get_gene_info'. No when-to-use or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_biological_claimVerify Biological ClaimCRead-onlyIdempotent
Verify a biological claim with structured claim decomposition and relation-specific database evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | Yes | Natural language biological claim to verify. | |
| context_gene | No | Optional gene symbol to focus evidence gathering. | |
| max_evidence_sources | No | Maximum evidence providers to query. Default 5. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, indicating safe, idempotent behavior. The description adds 'structured claim decomposition' but does not elaborate on query behavior, evidence sources, or result format. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise but uses jargon ('structured claim decomposition', 'relation-specific database evidence') that reduces clarity. The description could be more informative without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite good annotations and schema, the description lacks explanation of output format, how decomposition works, or how to interpret evidence. For a complex verification tool, this is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3). The description adds no information about the parameters (claim, context_gene, max_evidence_sources) beyond what is in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool verifies biological claims using structured decomposition and database evidence, which distinguishes it from sibling tools focused on specific analyses (e.g., pathway_analysis, variant_analysis). However, 'structured claim decomposition' and 'relation-specific database evidence' are somewhat vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., search_pubmed, get_gene_info). The description does not mention when not to use it or provide context for selecting it over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Tools cover distinct bioinformatics tasks with detailed descriptions, but some overlap exists (e.g., find_protein vs get_protein_info, pathway_analysis vs network_enrichment). Most tools have clear, differentiated purposes.
All tools use underscore_separated lowercase names following a verb_noun pattern (e.g., get_gene_info, search_pubmed). A few tools include version suffixes or omit verbs, but overall pattern is highly consistent.
32 tools is high for a single server, but the broad scope of bioinformatics (genes, proteins, drugs, variants, etc.) justifies the count. However, the number feels slightly heavy and could be streamlined.
The server covers a wide range of bioinformatics workflows including gene/protein lookup, variant analysis, drug repurposing, clinical trials, and network enrichment. Minor gaps exist (e.g., no GO enrichment), but overall coverage is comprehensive.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP gateway federating 21 biomedical MCP servers behind one endpoint: gnomAD, ClinVar, HPO, VEP.
Connect AI clients to biomedical data and tools.
MCP server for querying BrainKB, a knowledge base for neuroscience knowledge graphs.
Auditable MCP server for PubMed, Europe PMC, ClinicalTrials.gov, and bioRxiv/medRxiv queries
Related MCP Servers
- AlicenseAqualityCmaintenanceAn MCP server that provides standardized access to biomedical knowledge bases and resources, enabling AI systems to retrieve verified information from sources like bioRxiv, EuropePMC, and various protein/gene databases.5228Apache 2.0
- FlicenseNot gradedqualityDmaintenanceAn advanced integrated MCP server platform that combines 600+ tools and multiple biomedical databases to enable comprehensive information retrieval across molecules, proteins, genes, and diseases for accelerating therapeutic research.38
- AlicenseNot gradedqualityDmaintenanceA unified MCP server for biomedical research that connects AI systems to resources like Ensembl, EuropePMC, STRING, and more, enabling retrieval of verified domain-specific information.Apache 2.0
- AlicenseNot gradedqualityCmaintenanceUnified MCP server providing AI-agent-ready access to AlphaFold, PubMed, ChEMBL, Ensembl, and 37+ scientific databases.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/SachinGawande2003/Heuris-BioMCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server