Biomedical Research MCP Server
This server provides MCP tools for biomedical research, covering dilution calculations, experiment management, research document handling, PDF text extraction, and local evidence retrieval.
Calculate dilutions using C1V1 = C2V2 (stock, desired, final volume).
Retrieve experiment details by experiment ID.
Search experiments by query, cell type, treatment, organism, test, duration, and limit.
Add new experiments with metadata to the database.
Get research document metadata by document ID.
Add research document metadata (title, file name, description, type).
Extract PDF text from a document and store content/pages.
Retrieve extracted research content by document ID.
Search extracted content using SQL text matching (query, document, limit).
Search biomedical evidence via local retrieval system with question, top_k, and document filtering.
Provides tools for searching, retrieving, and adding biomedical experiment records stored in a local SQLite database.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Biomedical Research MCP ServerSearch experiments for doxorubicin toxicity in cardiac cells"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Biomedical AI MCP
A local Biomedical Research Assistant and Model Context Protocol (MCP) server for structured experiment data and evidence-based research-document retrieval.
Overview
This project combines a Python MCP server, a deterministic research agent, a SQLite database, PDF extraction, and local evidence retrieval.
The system is designed as a portfolio project demonstrating practical skills in:
Python application design
MCP server/tool development
Biomedical research data handling
SQLite database integration
PDF text extraction
Local evidence retrieval
Deterministic tool planning
Multi-tool execution and failure isolation
Evidence provenance and validation
Production-oriented logging
Automated testing
The current implementation is intentionally local and does not require a paid external LLM or paid embedding API.
Related MCP server: Researcher AI
Architecture
User
|
v
Biomedical Research Agent
|
+--> Planner
| |
| +--> selects MCP tools
|
+--> Executor
| |
| +--> executes tools
| +--> isolates failures
| +--> records execution status
| +--> logs execution events
|
+--> Validator
| |
| +--> checks cross-tool alignment
|
+--> Answer Generator
|
+--> evidence
+--> experiments
+--> document metadata
+--> provenance
|
v
MCP Server
|
+----------+-----------+
| | |
v v v
Research Experiment Document
Tools Tools Tools
| | |
+----------+-----------+
|
+-------+-------+
| |
v v
SQLite DB Research PDFs
|
v
PDF text extraction
|
v
Local evidence retrievalSee docs/architecture.md for the detailed architecture description.
Main Components
MCP Server
server.py exposes the biomedical capabilities through MCP.
Current MCP tools:
calculate_dilutionget_experimentsearch_experimentsadd_experimentget_research_documentadd_research_documentextract_pdf_textget_research_contentsearch_research_contentsearch_research_evidence
Current MCP resources:
research://experimentsresearch://experiments/{experiment_id}research://documentsresearch://documents/{document_id}research://documents/{document_id}/contentresearch://documents/{document_id}/pages/{page_number}
Agent
The agent is split into focused modules:
agent.py— compatibility entry pointagent_app.py— application flow and MCP connectionagent_planner.py— deterministic intent detection and tool planningagent_executor.py— MCP tool execution and failure isolationagent_validator.py— cross-tool validationagent_answer.py— research-answer formatting
Data Layer
database.py— SQLite database operationsbiomedical.db— project databasedocuments/— research PDFsretrieval.py— local evidence retrieval
Research Workflow
A research question follows this general path:
The user submits a biomedical research question.
The planner identifies the required tool domain.
The planner selects one or more MCP tools.
The executor calls the selected tools.
Individual tool failures are isolated so successful results can still be used.
Research evidence is retrieved from extracted document pages.
Experiment data can be retrieved from SQLite.
The validator checks relationships between retrieved evidence and database experiments.
The answer generator produces a structured response with evidence and provenance.
For combined questions, the planner can execute multiple domains in a deterministic order.
Example:
What mechanisms are involved in copper nanoparticle toxicity, and what related experiments are in the database?
This can result in:
search_research_evidence
search_experimentsThe answer can then distinguish literature evidence from database experiment metadata instead of treating them as the same source.
Local Evidence Retrieval
The project uses a local retrieval implementation in retrieval.py.
It:
loads extracted document pages
tokenizes and expands research questions
detects biomedical concepts and mechanisms
scores relevant pages
creates evidence snippets
returns page-level provenance
The current implementation is deliberately lightweight and deterministic. It does not depend on a paid external embedding service.
This makes the project reproducible in a local development environment, while also making its limitations explicit: the retrieval system is not equivalent to a large-scale neural semantic-search stack.
Example Evidence Domains
The current research workflow includes concepts such as:
oxidative stress
reactive oxygen species (ROS)
mitochondrial damage
lipid peroxidation
DNA damage
apoptosis
cell membrane damage
cell death and reduced viability
The answer layer preserves the distinction between direct and supporting evidence where available.
Database
The SQLite database contains structured experiment metadata.
The current experiment schema includes:
experiment_idnamecell_typetreatmentorganismtestduration_hours
The database layer uses parameterized SQL and validation rather than interpolating user-controlled values directly into SQL statements.
PDF Research Pipeline
Research documents are registered in the database and can be processed through the MCP server.
The extraction workflow is:
PDF
|
v
PDF metadata
|
v
pypdf text extraction
|
+--> full document content
|
+--> page-level content
|
v
SQLite storage
|
v
local evidence retrievalThe project includes a research PDF related to copper nanoparticle toxicity.
Error Handling
The agent is designed so that one failed tool does not necessarily invalidate the complete request.
Execution status is tracked separately:
_execution_status
|
+--> tool A: SUCCESS
|
+--> tool B: FAILEDThis allows the answer layer to distinguish complete and partial results.
MCP-level error results are also detected explicitly.
Observability
Logging is centralized through config.py.
The application records operational events such as:
tool-plan start and completion
individual tool start and completion
execution duration
successful tool execution
failed tool execution
MCP errors
unsupported tool requests
The executor intentionally avoids logging the complete research question.
Logging configuration is not duplicated across individual modules.
Configuration
Configuration is centralized in config.py and can be controlled through environment variables.
Important settings include:
BIOMED_DATABASE_PATHBIOMED_DOCUMENTS_DIRBIOMED_DEFAULT_DOCUMENT_IDBIOMED_RESEARCH_TOP_KBIOMED_MAX_QUESTION_LENGTHBIOMED_LOG_LEVELBIOMED_MCP_SERVER_COMMANDBIOMED_MCP_SERVER_SCRIPT
See .env.example for the supported configuration pattern.
Requirements
The current project targets:
Python 3.14
MCP Python SDK 2.1.1
pypdf 6.17.0
SQLite
uv (recommended for environment management)
The MCP server can be inspected with the MCP development tooling.
Setup
From the project directory:
uv syncOr, if the existing virtual environment is already configured:
.venv\Scripts\activateThe project currently uses a local SQLite database and local PDF files.
Run the MCP Server
The standard development command is:
uv run mcp dev server.pyThis opens the MCP development/inspection workflow and allows the exposed tools and resources to be inspected.
Run the Agent
From the project directory:
.venv\Scripts\python.exe agent.pyThe agent connects to the MCP server using the current Python interpreter.
Run Tests
Run the complete suite:
.venv\Scripts\python.exe run_all_tests.pyThe current verified baseline is:
Tests run: 126
Failures: 0
Errors: 0The test suite covers configuration, security, SQL safety, PDF limits, MCP behavior, planning, execution failure handling, answer quality, logging, and observability.
Example Questions
Research evidence
What mechanisms are involved in copper nanoparticle toxicity?Research + experiments
What mechanisms are involved in copper nanoparticle toxicity, and what related experiments are in the database?Experiment search
Which experiments use doxorubicin?Document information
Show me the research document.Combined document + evidence
Show me the paper and tell me what it says about apoptosis.Security and Safety Considerations
The project includes defensive controls around:
SQL query construction
input validation
numeric validation
document identifiers
PDF processing
tool limits
MCP errors
unsupported tools
execution failures
The system should still be treated as a research-support prototype rather than a clinical decision-support system.
It does not establish clinical efficacy, safety, diagnosis, or treatment recommendations.
Current Limitations
This is a portfolio/research prototype.
Known limitations include:
local deterministic retrieval rather than a production-scale vector database
limited document corpus
SQLite rather than a production database service
deterministic planning rather than a general-purpose LLM planner
no authentication layer for a deployed multi-user service
no production web UI
no clinical validation
evidence retrieval quality depends on the indexed document content
These limitations are intentional and documented rather than hidden.
Project Structure
biomedical-ai-mcp/
|
+-- agent.py
+-- agent_app.py
+-- agent_answer.py
+-- agent_executor.py
+-- agent_planner.py
+-- agent_validator.py
|
+-- server.py
+-- database.py
+-- retrieval.py
+-- config.py
|
+-- biomedical.db
+-- documents/
|
+-- tests/
+-- docs/
|
+-- .env.example
+-- .gitignore
+-- CONTRIBUTING.md
+-- LICENSE
+-- pyproject.toml
+-- README.md
+-- run_all_tests.pyLegacy and experimental files are kept separately under archive/legacy/ rather than mixed into the production-facing project structure.
Roadmap
The project roadmap includes:
production documentation
architecture documentation
Git/GitHub preparation
final end-to-end demonstration
stronger retrieval strategies
broader document collections
additional biomedical tools
optional LLM-based agent orchestration
deployment hardening
Portfolio Positioning
This project demonstrates an end-to-end biomedical AI engineering workflow rather than a single isolated model.
The strongest portfolio themes are:
MCP protocol implementation
biomedical research workflows
tool-using agent architecture
local retrieval/RAG concepts
structured data + unstructured documents
multi-tool orchestration
evidence provenance
defensive engineering
automated testing
production-oriented observability
License
See LICENSE.
Available Tools
10 toolsadd_experimentC
Add a new experiment to the database.
| Name | Required | Description | Default |
|---|---|---|---|
| test | No | ||
| organism | No | ||
| cell_type | No | ||
| treatment | No | ||
| experiment_id | Yes | ||
| duration_hours | No | ||
| experiment_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states only that an experiment is added, but does not mention duplicate handling, required concept relationships, side effects, idempotency, or any validation behavior. This is insufficient for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler or repetition. It communicates the core action and resource immediately, which is appropriate for a short purpose statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter creation tool with no output schema, no annotations, and no parameter documentation, this description is far too sparse. An agent cannot determine what data is needed, what the response will be, or how failures are handled, making the tool risky to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 parameters, and the description provides no parameter-level meaning beyond 'experiment'. Required versus optional fields, the role of fields like test, organism, cell_type, treatment, and duration_hours, and how they relate to the experiment are entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Add') and resource ('experiment to the database'), clearly indicating a creation operation. It distinguishes itself from sibling tools like get_experiment, search_experiments, and add_research_document by naming the experiment resource explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as add_research_document or update/search tools. The 'Add' phrasing implies creation, but there are no stated conditions, prerequisites, or exclusions to help an agent select it confidently.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_research_documentC
Add research document metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| file_name | Yes | ||
| description | No | ||
| document_id | Yes | ||
| document_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden of behavioral disclosure. It only says 'Add research document metadata' without explaining whether this creates a new record, updates an existing one, requires an existing experiment, or has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler or wordiness. It is front-loaded with the action and resource, though it is arguably too terse to fully support tool selection and invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0% parameter coverage, the description leaves the agent with many unknowns: what 'metadata' includes, what the document_id refers to, what document_type values are valid, and what happens after calling the tool. This is inadequate for correct invocation in non-obvious contexts.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention any parameters or their meanings. The tool has five parameters, three required, and the description adds no semantic value beyond the property names already present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Add') and resource ('research document metadata'), making the tool's core function clear. It does not explicitly differentiate from sibling tools like get_research_document or add_experiment, but 'document metadata' does imply a metadata-only scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives such as get_research_document or extract_pdf_text. No context is given about prerequisites, typical use cases, or situations where a sibling tool would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calculate_dilutionC
Calculate dilution using C1V1 = C2V2.
| Name | Required | Description | Default |
|---|---|---|---|
| stock | Yes | ||
| desired | Yes | ||
| final_volume | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations describing side effects, safety, or read-only/destructive nature. The verb 'calculate' implies a pure computation with no side effects, but this is not explicitly stated, so the agent cannot be certain that no data modifications occur. The description does not disclose any behavioral aspects beyond the calculation itself.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, consisting of a single sentence, which adheres to the principle of brevity. However, it lacks any structural organization (e.g., bullet points or sections) and is so terse that it omits critical details, making it less effective than a slightly longer but more structured description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is severely incomplete for an agent to use the tool correctly. It does not explain the formula's application, the meaning of each parameter, what output to expect, or any edge cases (e.g., division by zero). With no output schema and ambiguous parameters, an agent would be unable to determine the correct inputs or interpret the result, rendering the tool effectively unusable without external knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no descriptions for parameters, and the names 'stock', 'desired', and 'final_volume' are ambiguous in the context of the formula. It is unclear whether 'stock' refers to concentration or volume, and which parameter is the unknown being calculated. The lack of any explanation makes it impossible to correctly interpret the parameters without additional knowledge.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Calculate') and the subject ('dilution') with the formula C1V1 = C2V2, making the primary purpose evident. However, it does not specify what result is produced (e.g., the missing volume), so the output is not fully clear, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like search_experiments or add_research_document. The description gives no context about typical scenarios (e.g., preparing a solution) or how it relates to sibling tools, leaving the agent to infer its applicability.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_pdf_textB
Extract text from a PDF and save it into research_content and research_pages.
| Name | Required | Description | Default |
|---|---|---|---|
| document_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing side effects. It does state that the tool writes into two specific stores, which is helpful. However, it does not say whether the operation overwrites existing content, creates duplicates, requires special permissions, or behaves on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence states the action first and the storage target second. There is no redundant text or filler; every part of the sentence is informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a mutation tool with no annotations and no output schema, so the return/verification behavior is unspecified. It also lacks details on what happens if the document_id is invalid or the PDF cannot be parsed. The side-effecting behavior is identified, but an agent still cannot predict the result or how to confirm success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the sole required parameter document_id. The parameter is reasonably inferable from its name/title, but the description adds no value beyond the schema, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (extract text from a PDF) and the resulting side effect (saving into research_content and research_pages). It distinguishes itself from get/search/list siblings by describing an ingest operation, though it does not explicitly name sibling alternatives, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use case is implied: call this when a PDF's text needs to be extracted and persisted for later research. However, there is no explicit when-to-use guidance, no prerequisites, and no mention of when a sibling such as add_research_document would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_experimentB
Retrieve one experiment by experiment ID.
| Name | Required | Description | Default |
|---|---|---|---|
| experiment_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to fall back on, the description carries the full burden of disclosing behavior. It only states 'Retrieve', which implies a read-only operation, but does not mention any side effects, error conditions, or return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no redundant or extraneous information. It gets straight to the point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple retrieval operation but omits details that would make it fully complete, such as what the response looks like (since there is no output schema) or any potential errors. It is not severely lacking, but it leaves some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only provides the parameter name 'experiment_id' and its type. The description does not add any additional meaning about what constitutes a valid ID, how to obtain it, or any constraints, leaving the parameter semantics entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Retrieve') and the target ('one experiment'), with a specific identifier (experiment ID). It is unambiguous and directly conveys the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives like 'search_experiments'. It lacks explicit context about choosing this tool for direct ID-based lookup rather than searching.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_research_contentC
Retrieve extracted research document content.
| Name | Required | Description | Default |
|---|---|---|---|
| document_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It implies a read operation and that content is 'extracted,' but it does not explain return format, whether raw or structured content is returned, or any limitations. This is minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It is appropriately concise for a simple retrieval tool, though it lacks additional useful context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations, output schema, and usage guidance, the description is not complete enough for an agent to confidently select and invoke this tool. It does not clarify the relationship to get_research_document or search_research_content, nor what 'extracted' content means in practice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate for the single parameter, document_id. It does not describe what the document ID represents, how to obtain it, or any format expectations. The parameter name and title provide some self-evident meaning, but the description adds nothing beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Retrieve extracted research document content.' It is reasonably specific about the operation and object, though it does not explicitly differentiate from sibling tools like get_research_document or search_research_content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage context is provided. The description does not say when to use this tool instead of get_research_document, search_research_content, or extract_pdf_text, and there are no exclusions or alternative recommendations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_research_documentB
Retrieve research document metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| document_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. 'Retrieve research document metadata' implies a read-only metadata lookup, but it does not disclose what metadata fields are returned, behavior for missing documents, or error conditions. It is minimally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It states the action and object directly and is appropriately sized for a simple retrieval tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the tool is simple, the absence of annotations, output schema, and usage guidance makes the description incomplete. An agent does not know what metadata is returned, how to distinguish this from get_research_content, or how to handle potential failures.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for explaining the parameter. It does not mention document_id format, type, meaning, or any constraints. The parameter name and title are self-explanatory to a degree, but the description adds no clarification.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'Retrieve', with a clear resource, 'research document metadata.' It distinguishes itself from sibling get_research_content by specifying metadata rather than content, though it does not explicitly name that alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus siblings like get_research_content, search_research_content, or get_experiment. There is no mention of prerequisites, fallback tools, or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_experimentsC
Search experiments using:
general text query
cell type
treatment
organism
test
duration in hours
| Name | Required | Description | Default |
|---|---|---|---|
| test | No | ||
| limit | No | ||
| query | No | ||
| organism | No | ||
| cell_type | No | ||
| treatment | No | ||
| duration_hours | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only lists search dimensions and does not explain result behavior, whether filters are ANDed, whether query is a full-text search, pagination, default limits, or any side effects. This is a meaningful gap for a tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and structured as a scannable bullet list. Every line adds a distinct filter option and there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, no output schema, and no annotations, the description is incomplete. It lacks guidance on query semantics, filtering behavior, the purpose of the limit parameter, and how results are returned or sorted. An agent would need additional inference or trial-and-error to use the tool reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add value by listing six of the seven parameters and clarifying 'general text query' as the query parameter and 'duration in hours' as the duration filter. However, it omits the 'limit' parameter entirely and does not explain how multiple filters interact or what the default behavior is.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Search' and the resource 'experiments', and enumerates the filterable fields such as cell type, treatment, organism, test, and duration in hours. However, it does not explicitly differentiate itself from sibling search tools like search_research_content or search_research_evidence, so it is clear but lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to use this tool versus alternatives such as get_experiment or search_research_content. It does not mention whether filters combine, how to scope a search, or when the other search tools would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_research_contentC
Search extracted research pages using SQL text matching.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| document_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of behavioral disclosure. It reveals the underlying mechanism ('SQL text matching') but does not mention return behavior, pagination, result ordering, limitations, or whether the operation is read-only or has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loaded, with no filler or redundancy. The main action and target are stated immediately, though the brevity leaves substantial gaps in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three parameters, no output schema, and no annotations, this description is too sparse to fully guide an agent. The agent is left uncertain about what inputs mean, what the result will look like, and how this search differs from search_research_evidence or search_experiments.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the missing parameter documentation, but it does not. 'Search' implies the role of the query parameter, but limit and document_id are not explained at all, and the description adds no detail about how they affect the search.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search'), a clear resource ('extracted research pages'), and the method ('SQL text matching'). It is not a tautology and gives the agent a concrete idea of what the tool does, though it does not explicitly distinguish itself from siblings like search_research_evidence or search_experiments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus the sibling tools. The description simply states what it does, without saying when it is preferred, when it should be avoided, or what alternative should be used for different search needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_research_evidenceC
Retrieve relevant biomedical evidence using the local retrieval system.
| Name | Required | Description | Default |
|---|---|---|---|
| top_k | No | ||
| question | Yes | ||
| document_id | No | DOC001 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It only indicates a non-mutating local retrieval operation but does not disclose ranking behavior, return format, scope limits, pagination, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short sentence with no filler and leads with the action verb. It is efficient, though brevity comes at the cost of useful contextual information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With three parameters, zero schema descriptions, no output schema, and no annotations, the description is too thin for an agent to call the tool confidently. The ambiguity against similar sibling tools adds to the incompleteness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention question, top_k, or document_id. The agent is left without any explanation of the required argument or the defaults, beyond raw parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('retrieve') and a resource ('biomedical evidence') with a mechanism ('local retrieval system'). However, it does not distinguish this tool from sibling search_research_content or get_research_content, and 'evidence' remains somewhat loosely scoped.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as search_research_content or search_experiments. No usage conditions, exclusions, or selection criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
add_experiment - First observed
add_research_document - First observed
calculate_dilution - First observed
extract_pdf_text - First observed
get_experiment - First observed
get_research_content - First observed
get_research_document - First observed
search_experiments - First observed
search_research_content - First observed
search_research_evidence
TDQS
Scored across 10 tools
Most tools have clearly distinct purposes: experiment operations, document metadata, extracted content, evidence search, and calculation are separated by domain object. The only mild ambiguity is between search_research_content and search_research_evidence, but their descriptions indicate different search mechanisms.
All tool names follow a consistent snake_case verb_noun pattern using get, add, search, extract, and calculate. The naming is predictable and makes the action and target clear across the entire set.
Ten tools is well within the ideal range for a domain-specific server. Each tool covers a distinct aspect of experiment management, research document handling, content extraction, evidence retrieval, or calculation without unnecessary overlap.
The server supports creating, retrieving, and searching experiments and research documents, plus PDF content extraction and evidence lookup, which covers primary workflows. However, there are no update or delete operations for experiments or documents, and research document metadata cannot be searched directly, leaving notable lifecycle gaps.
Maintenance
Related MCP Connectors
Auditable MCP server for PubMed, Europe PMC, ClinicalTrials.gov, and bioRxiv/medRxiv queries
Bioinformatics MCP for genomic variant interpretation, gene-disease evidence and literature.
MCP gateway federating 22 biomedical MCP servers behind one endpoint: gnomAD, ClinVar, HPO, VEP.
PubMed MCP — wraps the NCBI E-utilities API (biomedical literature, free, no auth)
Related MCP Servers
- FlicenseCqualityDmaintenanceExposes a local biomedical literature pipeline as MCP tools for automated research workflows. Enables literature search, open-access paper retrieval, and draft generation for biomedical and pathology domains through standard MCP clients.6-
- AlicenseNot gradedqualityAmaintenanceEnables AI-assisted scientific research workflow management through MCP, including project creation, ideation, experiment execution, and artifact handling, with integration for ChatGPT, Codex, and Claude Code.Apache 2.0
- AlicenseNot gradedqualityAmaintenanceEnables MCP clients to interact with a local-first research knowledge workbench, supporting literature search, evidence-grounded Q&A, and reference export.2AGPL 3.0
- AlicenseBqualityAmaintenanceEnables provenance-first scholarly retrieval, paper ingestion, and reproducible research workflows by searching academic and developer sources, extracting source-located facts and claims, and preserving evidence and provider uncertainty for MCP clients.371Apache 2.0