biotools-mcp
# biotools-mcp
**Verified bioinformatics tools for AI agents** — sequence utilities + statistics, backed by
BioPython/scipy, exposed as an MCP server, companion Skill, and Python package.
[](https://pypi.org/project/biotools-mcp/)
[](https://pypi.org/project/biotools-mcp/)
[](LICENSE)
[](https://github.com/madhusudan-kulkarni/biotools-mcp/actions/workflows/ci.yml)
MIT · public · benchmarked.
> Research agents hallucinate bioinformatics math. This project builds the missing layer:
> battle-tested scientific computation wrapped in a clean, agent-native surface.
## What's inside
- **Sequence utilities** — GC content, reverse complement, translation, ORF finding, motif
scanning, sequence stats (BioPython-backed; ORF finder + motif-overlap are custom logic,
checked against independent references).
- **Statistics** — descriptive stats, t-test, chi-square, Mann–Whitney U, correlation
(scipy-backed, reference-vector tested against published/R values).
- **MCP server** — stdio + streamable HTTP, 11 action-oriented tools
(`seq_gc_content`, `stats_t_test`, …), pydantic v2 schemas with structured output.
- **Companion Skill** — agentskills.io spec: when to use which tool, input requirements,
and what **not** to do.
Every tool is reference-vector tested against published values (GenBank records, R `t.test`
output, Mendel's 1866 pea counts, Anscombe's quartet) — **never** against the wrapper itself.
## Install
Published on PyPI as `biotools-mcp` (requires Python 3.11+).
```bash
# Install as a standalone CLI (MCP server entry point)
uv tool install biotools-mcp
# Or run without installing
uvx biotools-mcp
# Or plain pip
pip install biotools-mcp
```
As a library:
```bash
uv add biotools-mcp
```
```python
from biotools_mcp.seq import gc_content
from biotools_mcp.stats import t_test
print(gc_content("ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG").gc_percent) # 56.4103
print(t_test([1, 2, 3], [4, 5, 6]).statistic) # -3.6742
```
## MCP configuration
Point your agent at the server over stdio:
**Claude Code** (`.mcp.json`):
```json
{
"mcpServers": {
"biotools-mcp": {
"command": "uvx",
"args": ["biotools-mcp"]
}
}
}
```
**Cursor** — Settings → MCP → Add:
```json
{
"mcpServers": {
"biotools-mcp": {
"command": "uvx",
"args": ["biotools-mcp"]
}
}
}
```
**Codex / Gemini CLI** — same `mcpServers` block in the agent's MCP config file.
**Streamable HTTP** (for remote use):
```bash
uvx biotools-mcp --transport streamable-http
```
## Tools
| Tool | Description |
|---|---|
| `seq_gc_content` | GC content as a percentage (ambiguous bases excluded) |
| `seq_reverse_complement` | Reverse complement (IUPAC-aware, DNA/RNA) |
| `seq_translate` | Translate to protein (NCBI tables, incl. mitochondrial) |
| `seq_orf_finder` | Open reading frames, forward strand, frames 0–2 |
| `seq_motif_scan` | IUPAC motif scanning with bracket groups + overlap policy |
| `seq_stats` | Length, mono/di composition, GC skew |
| `stats_describe` | Descriptive statistics (n, mean, median, var, skew, kurtosis) |
| `stats_t_test` | Student/Welch two-sample t-test with Cohen's d |
| `stats_chi_square` | Chi-square test of independence (Yates optional) |
| `stats_mann_whitney` | Mann-Whitney U test (exact/asymptotic) |
| `stats_correlation` | Pearson or Spearman correlation |
## Benchmarks
Wedge subset of BioAgent Bench + BioTaskBench (sequence utilities + statistics) run against
the tools and published in [benchmarks/results.md](benchmarks/results.md). Every task in the
subset runs — failures are published alongside passes.
**Current: 15/15 passed (100%)** — bioagent-bench subset 7/7, bioTaskBench subset 8/8.
Pinned harness, see the results file for task-level detail.
```bash
uv run python benchmarks/run_bioagent_bench.py # 7/7
uv run python benchmarks/run_biotaskbench.py # 8/8
uv run python benchmarks/harness.py --suite all # combined 15/15
```
## Companion Skill
The [skills/biotools-mcp](skills/biotools-mcp/SKILL.md) skill (agentskills.io spec) teaches
agents the tool inventory, when to use which tool, input requirements, and what **not** to do
(never compute GC/translation/t-tests by hand). Load it into any skills-compatible agent.
## Documented solutions
Past problems and the patterns they produced live in
[docs/solutions](docs/solutions/README.md) — including the `uvx` grandchild-process leak in
subprocess tests, mcp SDK v2 tool-registration conventions, and the CI matrix Python-version
trap. Relevant when implementing or debugging in those areas.
## Development
```bash
uv sync --extra dev
uv run pytest # reference-vector suite (skips slow packaging by default)
uv run python benchmarks/harness.py --suite all # regenerate benchmark results
uvx ruff check src tests benchmarks
```
CI runs lint + tests on Python 3.11/3.12; a nightly workflow regenerates the benchmark table;
a tag-pushed `v*` triggers the PyPI publish workflow.
## Contributing
See [CONTRIBUTING.md](CONTRIBUTING.md) — verification is non-circular
(reference-vector fixtures from published values), degenerate inputs must
fail loudly, and benchmark tasks are never dropped to keep numbers green.
Changes are tracked in [CHANGELOG.md](CHANGELOG.md).
## License
MIT — see [LICENSE](LICENSE).
TDQS
Scored across 11 tools
Each tool has a distinct purpose within its domain. The seq_* tools cover motif scanning, GC content, reverse complement, translation, ORF finding, and sequence stats with no overlap. The stats_* tools cover descriptive statistics, t-test, chi-square, Mann-Whitney, and correlation, also without ambiguity. The two domains are clearly separated by prefix and description.
Tool names follow a consistent domain-prefix pattern: seq_ for sequence operations and stats_ for statistical tests. However, within each prefix, the naming style is mixed (e.g., seq_translate is a verb, seq_orf_finder is a noun; stats_describe is a verb, stats_t_test is a noun). This minor inconsistency lowers the score from 5 to 4, but the prefix convention makes names predictable.
With 11 tools, the server is well-scoped for a bioinformatics toolkit. It covers a reasonable set of sequence analysis functions and common statistical tests without being bloated. The count falls well within the ideal range for a focused utility server.
The tool surface covers core sequence operations (translation, reverse complement, GC content, motif scanning, ORF finding, and stats) and common statistical tests (descriptive, t-test, chi-square, Mann-Whitney, correlation). Minor gaps exist, such as sequence alignment or advanced statistical tests like ANOVA, but agents can perform most basic workflows without dead ends.