BioHarbor
BioHarbor
Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.
π§ͺ Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome β see the roadmap.
BioHarbor is an MCP server that lets AI agents such as Claude, Cursor and Codex execute bioinformatics tools β not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.

Why another bio MCP server?
Most bio MCP servers wrap databases (UniProt, PDB, PubMedβ¦). Use them β BioHarbor complements them by running the compute:
Database MCP servers | BioHarbor | |
Runs real analyses (search, fold, cluster) | β | β |
Validates inputs before burning GPU time | β | β |
GPU-aware queue, polite on shared GPUs | β | β |
Long jobs return a | β | β |
Compact summaries + files on disk (saves tokens) | β | β |
Provenance for every run, export to a pipeline | β | β (export: planned) |
Related MCP server: BioMCP
Quick start
pip install bioharbor
bioharbor doctor # checks Python, GPUs, workspace, tools
bioharbor setup-db swissprot # reference database for search_homologs (needs MMseqs2)
bioharbor install # shows how to connect Claude, Cursor or CodexFor structure prediction on a GPU: pip install "bioharbor[esmfold]" β see
docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).
Connect your agent
BioHarbor is a standard MCP server, so it works with any MCP client. One command sets up the popular ones (it writes an absolute path, so GUI apps find it even outside your venv):
Client | Set up |
Claude Code |
|
Claude Desktop |
|
Cursor |
|
Codex (CLI, IDE extension, app) |
|
Anything else | run |
Long-running tools return a job_id within ~20 s instead of blocking, so they stay
within every client's tool-call timeout.
Step-by-step setup (local or on a GPU server, with troubleshooting): docs/connect-clients.md.
Then ask your agent something like:
Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?
Use it without an agent
Every tool is also a CLI command, with identical behaviour:
bioharbor tools list
bioharbor run find_orfs sequence=@contig.fa min_aa=100 --brief # human-readable
bioharbor run seq_stats sequence=MKTAYIAKQRQISFVKSHFSRQ
bioharbor jobsShared GPU server
GPUs on a lab server, agent on your laptop? Run BioHarbor on the server and reach it through an SSH tunnel; no extra port is opened on the server:
# on the GPU server
bioharbor serve --http --host 127.0.0.1 --port 8765
# on your laptop, then point Cursor / Codex / Claude Code at http://127.0.0.1:8765/mcp
ssh -N -L 8765:127.0.0.1:8765 you@gpu-serverβ οΈ HTTP mode has no authentication yet (on the roadmap), so keep it on
127.0.0.1and use the tunnel. Details: docs/connect-clients.md.
Tools
Tool | What it does | Runs |
| Validate sequences; type, length, GC%, molecular weight | inline |
| DNA/RNA β protein, one or all six frames | inline |
| Longest ORFs on both strands, with coordinates | inline |
| MMseqs2 search (protein, or translated DNA) vs local DBs | job |
| ESMFold structure, pLDDT bands, low-confidence regions, pTM | job (GPU) |
| scanpy QC β clustering β markers | planned |
Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases,
gpu_status, read_file.
How it works
Agent ββMCPβββΆ validate input ββΆ inline? ββyesβββΆ run ββ
β no βββΆ provenance + summary ββΆ Agent
βΌ β
job queue (SQLite) ββΆ GPU placement β
(waits politely for a GPU with free memory)Every call is a job recorded in SQLite with params, versions, timings and GPU used, plus a
provenance.jsonnext to its outputs.GPU placement reads live free memory and utilisation (NVML or
nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected.Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
Results are agent-shaped:
summary,message,files,suggestions. Errors carry ahintand aretryableflag.
Details: docs/design.md.
Writing a tool
from pydantic import BaseModel, Field
from bioharbor.registry import Resources, RunContext, tool
from bioharbor.results import ToolResult
class FoldParams(BaseModel):
sequence: str = Field(..., description="Protein sequence")
@tool(
version="1",
slow=True,
resources=Resources(gpu=True, gpu_mem_gb=lambda p: 4 + len(p.sequence) / 100),
)
def predict_structure(params: FoldParams, ctx: RunContext) -> ToolResult:
"""Predict a protein structure with ESMFold."""
...
return ToolResult(summary={"mean_plddt": 87.1}, files=["model.pdb"])Plugins can ship tools in their own package via the bioharbor.tools entry-point group.
See CONTRIBUTING.md.
License
Available Tools
12 toolscancel_jobB
Cancel a job that is still queued or waiting for a GPU.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses one useful constraint (only queued/waiting jobs are cancellable) but says nothing about permissions required, whether cancellation is reversible, what happens to a running job, or how errors surface for an invalid state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence stating the action and its constraint, with no filler. Nothing needs to be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the tool is trivially simple with one required parameter. Still, as an unannotated mutation tool it should at minimum state the consequence of cancelling and the failure mode for non-queued jobs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter job_id has 0% schema description coverage, so the description is the only place semantics could be added, yet job_id is never mentioned. No format, provenance (e.g., obtain via get_job/list_jobs), or validity rules are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Cancel a job') plus a scope qualifier (queued or waiting for a GPU), which clearly separates it from siblings like get_job, list_jobs and gpu_status. It does not name a sibling explicitly, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'still queued or waiting for a GPU' gives an implicit applicability condition, so the agent can infer this is for not-yet-running jobs. However, there is no guidance on what happens if the job is already running, no idempotency note, and no alternatives suggested.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_toolC
Full input schema, version and resource needs of a BioHarbor tool.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full disclosure burden. It never states that this is a non-destructive read, that it takes no job/session state, or whether it fails for unknown tool names. It partially compensates by listing what is returned, but the output schema already covers that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the returned content with zero filler. It errs only in being too sparse rather than too long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the tool is a simple single-parameter read. Still, the meaning of the required 'name' argument and the when-to-use context are left entirely unaddressed, leaving the definition minimally viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter is 'name' with 0% schema description coverage and no enum. The description does not clarify whether this is a tool name, job name, or resource name, nor the expected format or error behavior for unknown names. It fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (describe/extract metadata) and enumerates exactly what is returned: full input schema, version, and resource needs of a BioHarbor tool. That is far more informative than the name alone. It doesn't, however, distinguish itself from sibling tools explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives, nor any precondition such as 'call before invoking a tool to learn its arguments.' The use case (introspecting a tool before use) is only inferable from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_orfsB
Find open reading frames (ATG to stop) on both strands.
Returns the longest ORFs in DNA/RNA with coordinates (1-based, on the forward strand) and writes all ORFs to a FASTA file.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | How many of the longest ORFs to return. | |
| min_aa | No | Minimum ORF length in amino acids (1 = any ORF). | |
| sequence | Yes | A raw sequence or FASTA text (one or more records). | |
| both_strands | No | Also search the reverse complement. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It does disclose a side effectβwriting all ORFs to a FASTA fileβand mentions returned coordinates. However, it omits whether the file write overwrites existing files, required permissions, runtime characteristics, or whether the operation is otherwise read-only. Partial behavioral coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose. The second sentence efficiently covers return format and side effect. No redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained, though the description does so anyway. Missing usage guidelines and partial behavioral disclosure (e.g., file overwrite behavior) leave the definition incomplete for a tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description adds little beyond the schemaβ'both strands' and 'longest ORFs' map directly to parameters but without extra syntax or format details. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Find open reading frames (ATG to stop) on both strands.' Clear what the tool does, but does not explicitly distinguish from siblings like translate_sequence or seq_stats.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives named, no prerequisites mentioned. The description only implies usage through the purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_jobB
Status and (when finished) result of a background job.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job id returned by a slow tool. | |
| wait_seconds | No | Block up to this long for the job to finish. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully reveals the dual return behavior (status while running, result once finished), but says nothing about polling expectations, blocking semantics (wait_seconds is only explained in the schema), authentication, or what a failure/error state looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the primary purpose (status) front-loaded before the conditional result case. Nothing is wasted, though the extreme brevity leaves room that could have been spent on usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and both parameters are documented in the schema. What is missing is the calling context: how this fits with the slow tools that produce job ids and with list_jobs/cancel_job as alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so job_id and wait_seconds are already fully documented in the schema, including the blocking cap of 25 seconds. The description adds no parameter-level meaning beyond that, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (background job) and what it returns (status, and result once finished). The 'get' verb plus 'job' clearly separates it from list_jobs and cancel_job by naming convention alone, though the description itself never explicitly contrasts the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to call this versus list_jobs (to enumerate jobs) or cancel_job (to abort one). The only usage hint is carried by the schema ('Job id returned by a slow tool'), not by the description, so the agent must infer the polling workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gpu_statusA
Live GPU memory and utilisation (includes other users' processes).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses that output includes other users' processes, implying cluster-wide visibility rather than just the caller's jobs, which is genuinely valuable behavioral context. However, it says nothing about permissions, refresh/caching behavior, or whether output is a snapshot.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero waste. The most important qualifier ('live') and the scope note (other users' processes) are both packed into one line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not required here, and there are no parameters. The description is nearly complete for a zero-arg read tool; the only minor gap is the absence of any usage framing relative to the job-submission siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document; baseline 4 applies. The description correctly adds no parameter discussion, consistent with an empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (GPU memory and utilisation) with a clear scope qualifier ('live'), which an agent can distinguish from the sibling tools (seq_stats, list_jobs, etc.). It doesn't need to name a sibling because none of the others are GPU-related, but it also doesn't explicitly say it is read-only status retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided. The agent must infer that this is used to check GPU availability before submitting jobs (e.g., before predict_structure or search_homologs). There is no mention of alternatives, prerequisites, or timing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_databasesB
Sequence databases installed for search_homologs, and ones that can be added.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It states what the list contains (installed vs addable) but does not disclose that this is a read-only operation, nor any authentication, side-effect, or rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise phrase with no redundant content, and the key distinction (installed for search_homologs) is front-loaded. It is a fragment rather than a complete sentence, but it wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with an output schema, the description covers what is listed. However, it omits when to invoke it and does not state its read-only nature; with no annotations, more behavioral context would be useful, though the output schema handles return structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema description coverage is 100%, so the baseline is 4. The description adds no parameter information because there are none to describe.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (sequence databases) and distinguishes two subsets: installed for search_homologs and addable ones. While it lacks an explicit verb like 'lists', the tool name supplies that, and the mention of the sibling search_homologs helps an agent understand the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is provided. The reference to search_homologs implies these databases are meant for that tool, giving an implied usage context, but the description does not state that you should call this before search_homologs or how it differs from other siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_jobsA
Recent jobs, newest first (results omitted; use get_job for details).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No | Filter: queued, waiting_gpu, running, succeeded, failed, cancelled |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that results are omitted from the payload and that ordering is newest-first, but says nothing about read-only safety, pagination, or how 'recent' is bounded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the ordering and the key caveat front-loaded and zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description still helpfully notes results are omitted. The remaining gap is the undocumented limit parameter, a minor omission for such a simple listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: status is documented in the schema, but limit has no description anywhere. The description adds no meaning for either parameter, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource (recent jobs) and the ordering contract (newest first), and explicitly distinguishes itself from get_job, which is a named sibling. It is not a full verb+resource sentence, but the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Routes the agent explicitly: for details of a specific job, use get_job. That covers the main decision point, though it does not say when listing is preferable to other approaches or mention that no filtering beyond status is available.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_structureA
Predict 3D protein structure with ESMFold on a GPU.
Returns per-protein mean pLDDT, confidence bands, low-confidence regions and pTM; PDB files are written to disk (B-factor column = pLDDT).
Runs as a background job; may return a job_id to poll with get_job.
| Name | Required | Description | Default |
|---|---|---|---|
| sequence | Yes | Protein sequence(s), raw or FASTA (β€20 records). | |
| num_recycles | No | Recycling iterations; more can help hard targets. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a good job: it discloses asynchronous execution ('background job', 'may return a job_id to poll with get_job'), the computed confidence outputs, and the side effect of writing PDB files to disk with pLDDT in the B-factor column. It stops short of GPU contention/time expectations, but the operational profile is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: capability first, return artifacts second, execution model third. Front-loaded with the core action and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though an output schema exists, the description usefully complements it with output artifacts and async semantics, and it covers the GPU/background-job behavior an agent needs to call correctly. Only minor gaps (expected runtime, GPU capacity limits, failure behavior) remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both 'sequence' (raw or FASTA, β€20 records) and 'num_recycles' (more helps hard targets) documented in the schema itself. The description adds no parameter-level detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Predict 3D protein structure') plus the concrete method ('with ESMFold on a GPU'), which cleanly separates it from siblings like search_homologs, translate_sequence, or seq_stats. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage ('predict 3D protein structure') and adds operational routing ('Runs as a background job; may return a job_id to poll with get_job'), but never states when to prefer this tool over alternatives such as search_homologs for structural inference by homology. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileB
Read (the start of) a text output file produced by a BioHarbor tool.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | A path from a tool result's `files` list. | |
| max_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the read is partial ('the start of') and restricted to text output files, but it does not explain the max_bytes limit, truncation behavior, permission requirements, or error cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is a single, front-loaded sentence with no redundant or vague filler. It communicates the core operation and its partial-read nature immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, with no annotations and incomplete parameter documentation, the description leaves gaps around max_bytes semantics and exact usage context that an agent would need when invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: path is documented in the schema, but max_bytes is not. The description's '(the start of)' hints at byte limits but does not explain max_bytes or add any path semantics beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (read) and resource (text output file produced by a BioHarbor tool), and clarifies that only the start of the file is read. It distinguishes the tool from the computation-oriented siblings, which do not read output files. It stops short of a 5 because it does not name or contrast any alternative file-reading path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'produced by a BioHarbor tool' implies the tool is used on files returned in tool results, but there is no explicit when-to-use guidance, prerequisites, or alternative tools. Usage is therefore only implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_homologsA
Find homologous proteins with MMseqs2 (Swiss-Prot, PDB, ...).
Searches a local database (default Swiss-Prot). Returns the best hits per query with identity, E-value, coverage, organism and description; the full hit table is written to a TSV file.
Runs as a background job; may return a job_id to poll with get_job.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | Hits per query shown inline. | |
| evalue | No | E-value cutoff. | |
| database | No | Installed database name (see list_databases). | swissprot |
| max_hits | No | Max hits per query kept in the output file. | |
| sequence | Yes | Protein (or DNA, searched translated) sequence(s), raw or FASTA. | |
| sensitivity | No | MMseqs2 -s: 1 fast β¦ 7.5 most sensitive. | |
| min_coverage | No | Minimum alignment coverage of query and target. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the search source (local DB, default Swiss-Prot), the inline return shape (best hits with identity, E-value, coverage, organism, description), the side effect of writing a full hit table to a TSV file, and the async job semantics with polling via get_job. It omits any note on permissions, resource cost, or runtime expectations for a sensitive search.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short, front-loaded lines: purpose first, then return shape, then async behavior. No filler, and the job_id/get_job caveat is placed where an agent will see it before deciding to call.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, yet the description still usefully notes the TSV side output and the job_id polling path. It is essentially complete for invocation; only multi-sequence batching semantics and resource expectations are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters are already documented (top_n, evalue, max_hits, database, sequence, sensitivity, min_coverage). The description adds only loose framing ('best hits per query', the returned fields) rather than explaining interactions such as top_n vs max_hits or how sensitivity/min_coverage affect results. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Find homologous proteins') plus the underlying engine (MMseqs2) and database scope (Swiss-Prot, PDB, ...). This plainly distinguishes it from siblings like seq_stats, translate_sequence, and find_orfs, which do not perform homology searches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: it searches a *local* database, defaults to Swiss-Prot, and runs as a background job that may return a job_id to poll with get_job. The get_job routing is exactly the alternative-selection guidance an agent needs. It doesn't state when-not to use it or explicitly point at list_databases for valid DB names in the description body.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seq_statsA
Validate sequences: type, length, GC%, molecular weight.
Reports dna/rna/protein type, length, GC content (nucleotides) or molecular weight (proteins) and composition. A cheap first step before heavier tools.
| Name | Required | Description | Default |
|---|---|---|---|
| sequence | Yes | A raw sequence or FASTA text (one or more records). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the output content and cost profile ('cheap first step'), which is useful, but says nothing about read-only semantics, behavior on invalid input, or whether multiple FASTA records are handled. Adequate but incomplete for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the metric list. Slight redundancy between the first line and the restatement of what is reported keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the description still summarizes them. For a single-parameter read-only analysis tool, the coverage is essentially complete aside from invalid-input behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, which already explains it accepts a raw sequence or multi-record FASTA text. The description adds no format or constraint detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (validate) and resource (sequences) plus the metrics computed (type, length, GC%, MW), so an agent can tell it apart from translate_sequence or find_orfs. It does not, however, explicitly contrast itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing line 'a cheap first step before heavier tools' gives clear context on when to reach for this tool rather than the heavier siblings. No explicit exclusions or named alternatives are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
translate_sequenceA
Translate DNA/RNA to protein, in one frame or all six.
Uses the standard genetic code; stop codons appear as '*'.
| Name | Required | Description | Default |
|---|---|---|---|
| frame | No | Reading frame; negative = reverse complement; 'all' = six frames. | |
| to_stop | No | Stop at the first stop codon. | |
| sequence | Yes | A raw sequence or FASTA text (one or more records). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does add real context: the standard genetic code is used and stop codons render as '*'. However, it omits what happens with ambiguous bases, non-multiple-of-3 input, or invalid characters, which matters for a sequence-translation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core operation followed by the key output convention. Every sentence carries information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return formatting need not be explained, and the description covers the translation semantics and stop-codon convention. The only gap is behavior on multi-record FASTA input and malformed sequences, which is minor given the complete schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents frame (including negative/reverse-complement and 'all'), to_stop, and the FASTA-capable sequence input. The description only loosely echoes the frame behavior ('one frame or all six'), adding little beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Translate DNA/RNA to protein') plus the scope of frames, so an agent knows exactly what the tool produces. It does not distinguish itself from the sibling find_orfs, which also translates nucleotide sequence, so the sibling-differentiation credit is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in one frame or all six' implies when a given frame mode is appropriate, but there is no explicit guidance on when to use this tool rather than find_orfs (which surfaces ORFs) or seq_stats. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.1.0- First observed
cancel_job - First observed
describe_tool - First observed
find_orfs - First observed
get_job - First observed
gpu_status - First observed
list_databases - First observed
list_jobs - First observed
predict_structure - First observed
read_file - First observed
search_homologs - First observed
seq_stats - First observed
translate_sequence
TDQS
Scored across 12 tools
Each tool targets a distinct operation: sequence statistics, translation, ORF finding, homology search, structure prediction, and job management are clearly separated. No overlapping purposes; get_job, list_jobs, and cancel_job are distinct job operations.
Names uniformly use snake_case and mostly follow a verb_noun pattern (e.g., translate_sequence, find_orfs, get_job). Minor deviation: seq_stats and gpu_status are noun phrases, but overall consistent and readable.
12 tools is well within the ideal 3-15 range, each tool serves a clear purpose in the bioinformatics workflow without redundancy. The set is well-scoped for the domain.
Covers a sequence analysis pipeline from validation to structure prediction, with job management and resource introspection. Minor gap: no explicit tool for retrieving or loading input sequences (e.g., from files or databases), though read_file exists for outputs.
Maintenance
Related MCP Connectors
Protein analysis: ESM-2/ESMC embeddings, mutation scoring, landscape scans, ESMFold structure.
- mcpOAuthio.scispot
Turn any LLM into your lab assistant: search samples, track experiments, analyze data with AI.
AI orchestration for computational chemistry and HPC workflows.
MD workbench for agents: 13 analysis tools, renders, hosted GPU runs; LAMMPS + OpenMM.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables AI-powered protein design, including binders, peptides, and custom proteins, with GPU-accelerated Docker deployment and async job management.2-
- AlicenseAqualityAmaintenanceA high-performance MCP server that gives LLMs access to 25 biomedical tools federated across 50+ upstream APIs for genes, variants, drugs, diseases, literature, clinical trials, and structural biology.41205 npm12Apache 2.0
- AlicenseNot gradedqualityCmaintenanceUnified MCP server providing AI-agent-ready access to AlphaFold, PubMed, ChEMBL, Ensembl, and 37+ scientific databases.MIT
- AlicenseNot gradedqualityCmaintenanceMCP-native scientific skills for reproducible computational biology and AI-driven drug-discovery workflows. It combines deterministic scientific tools with an MCP server to give AI agents real computational capabilities.Apache 2.0