Skip to main content
Glama
nusduck

biz-db-mcp

by nusduck

biz-db-mcp

一个面向 Agent 的 MySQL MCP Server,支持多数据库、只读查询、受保护的 DML 写入,以及批量导出到 Sandbox Dataset。

v0.2 灵活性升级:按库覆盖配置、WHERE 白名单、list_tables/explain 勘探、dry_run 预检、execute_many 批量事务、offset 分页与 csv/jsonl 导出。

设计边界

  • 默认只读:BIZDB_WRITE_ENABLED 默认关闭,单个数据库还必须显式 allow_writes: true

  • queryexport 只允许单条 SELECT/WITH ... SELECTexecute 只允许单条 INSERTUPDATEDELETEREPLACEexecute_many 允许单事务批量(最多 100 条)。

  • UPDATE/DELETE 默认必须带 WHERE,可通过 BIZDB_WRITE_REQUIRE_WHERE_EXCEPT_TABLES 按表名 glob 放宽(如 temp_*,cache_*);写入在事务中执行,并受最大影响行数限制(支持按库覆盖)。

  • export 使用服务端游标分批写 Parquet/CSV/JSONL,再上传 Sandbox;身份从进程环境注入,模型不能指定目标 session。

  • SQL 审计只记录 SHA-256 和长度,不记录原始 SQL、参数或密码。

  • 所有限流与超时均支持按库覆盖,全局配置为 fallback。

Related MCP server: MySQL ReadOnly MCP Server

目录结构

biz_db_mcp/
  config.py          # 环境变量、多数据库、写入和 DBPM 配置(含按库覆盖与文件加载)
  db.py              # MySQL 连接、查询、写入、勘探(describe/list_tables/explain/批量)
  dbpm.py            # DBPM 外部桥接命令适配器
  sql_guard.py       # SELECT/DML 语句边界
  export.py          # 流式 Parquet/CSV/JSONL 导出
  sandbox_client.py  # Dataset 上传(多格式 MIME)
  audit.py           # 结构化审计日志
  server.py          # MCP 工具注册(8 个工具)
tests/               # 单元测试

配置

单库兼容配置:BIZDB_MYSQL_DSN=mysql://user:password@host:3306/database

多库使用 BIZDB_DATABASES_JSON,格式可以是映射或 connections 数组:

{"connections":[
  {"id":"primary","dsn":"mysql://user:password@db-a/app","query_max_rows":200},
  {"id":"reporting","dsn":"mysql://user:password@db-b/reporting","allow_writes":false,"query_max_rows":5000}
]}

也支持文件加载:BIZDB_DATABASES_FILE=/etc/biz-db.json(内容与 BIZDB_DATABASES_JSON 同格式,优先级:BIZDB_DATABASES_JSON > BIZDB_DATABASES_FILE > BIZDB_MYSQL_DSN)。

BIZDB_DEFAULT_DATABASE 选择默认库。写入还需要 BIZDB_WRITE_ENABLED=true,并在目标库配置 allow_writes: true;可用 BIZDB_WRITE_MAX_AFFECTED_ROWSBIZDB_WRITE_REQUIRE_WHERE 调整保护阈值。按库覆盖示例:

{"id":"writer","dsn":"mysql://...","allow_writes":true,"write_max_affected_rows":5000,"write_require_where":false,"query_timeout_seconds":30}

支持的按库覆盖字段:query_max_rowsquery_timeout_secondsexport_max_rowsexport_timeout_secondswrite_max_affected_rowswrite_require_wherenull 表示继承全局)。

WHERE 白名单:BIZDB_WRITE_REQUIRE_WHERE_EXCEPT_TABLES=temp_*,cache_logs(逗号或空格分隔,glob 匹配,大小写不敏感),命中白名单的表允许 UPDATE/DELETE 不带 WHERE

工具一览(8 个)

工具

说明

灵活性增强

list_databases

列出已配置库,不暴露 DSN

按库覆盖后仍脱敏返回

describe_table

返回表结构,支持 schema.table

支持 db.table 限定

list_tables

列出库内表/视图,支持 pattern(LIKE)过滤

新增

query

受控只读 SELECT,limit 受按库上限约束

新增 offset 分页、dry_run 预检

explain

返回 EXPLAIN 执行计划,不实际执行

新增

execute

单条 DML 事务提交

新增 dry_run(EXPLAIN/ROLLBACK 预估)

execute_many

多条 DML 单事务批量(≤100 条)

新增,atomic 控制是否全回滚

export

流式导出到 Sandbox Dataset

新增 format: parquet/csv/jsonl

execute 仍为单语句 DML 接口,不提供 DDL、事务控制或多语句拼接;批量场景请用 execute_many

DBPM(可选)

UPspec 中的 DBPM 是密码管理系统,不是 MCP 请求认证。Python 项目不内置私有 Java SDK;启用 BIZDB_DBPM_ENABLED=true 后,必须配置 BIZDB_DBPM_COMMANDBIZDB_DBPM_SERVERSBIZDB_DBPM_APP_NAMEBIZDB_DBPM_APP_KEYBIZDB_DBPM_APP_IP

桥接命令从 stdin 读取 JSON(包含 database、username、servers 等字段),只在 stdout 输出密码,非零退出码表示失败。DBPM 私钥通过 DBPM_APP_KEY 环境变量传给桥接进程,不出现在命令行参数中。数据库条目可用 dbpm_database/dbpm_user 覆盖 DBPM 查询键。

开发与运行

python3 -m pip install -e '.[dev]'
python3 -m pytest
biz-db-mcp

Sandbox 上传仍需配置 BIZDB_SANDBOX_BASE_URLBIZDB_SANDBOX_API_TOKEN;导出身份需配置 BIZDB_SESSION_IDBIZDB_ORG_IDBIZDB_USER_ID。生产环境应使用 DBPM,不要把真实密码提交到配置、日志或测试夹具中。

灵活性使用示例

# 分页查询
query("SELECT id, name FROM users WHERE active=1", limit=50, offset=100)

# 预检(不查库)
query("SELECT * FROM orders WHERE status=%s", ["paid"], dry_run=True)
# -> {"dry_run": True, "effective_limit": 200, "effective_timeout": 10}

explain("SELECT * FROM orders WHERE status=%s", ["paid"])
list_tables(pattern="order%", include_views=False)
describe_table("reporting.orders")

# 写入预估

execute("UPDATE users SET active=0 WHERE id=%s", [123], dry_run=True)
# -> {"dry_run": True, "estimated_plan": [...] }

# 批量事务
execute_many([
  {"sql": "INSERT INTO users(name) VALUES (%s)", "params": ["alice"]},
  {"sql": "UPDATE users SET active=1 WHERE id=%s", "params": [123]}
], atomic=True)

# 多格式导出
export("SELECT * FROM orders", dataset_name="orders_2024", format="csv")
export("SELECT * FROM orders", dataset_name="orders_2024", format="jsonl")

环境变量速查(新增标记 ★)

变量

说明

默认

BIZDB_DATABASES_FILE

多库 JSON 文件路径

-

BIZDB_WRITE_REQUIRE_WHERE_EXCEPT_TABLES

免 WHERE 白名单(glob)

-

单库按库覆盖 ★

query_max_rows 等 6 个字段可在 BIZDB_DATABASES_JSON 单库内覆盖

继承全局

Available Tools

5 tools
describe_tableB

Return column metadata for a table in one configured database.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes
databaseNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must disclose behavioral traits. It only mentions the return of metadata and a vague scope ('one configured database'), without addressing read-only status, error handling, or permissions. The phrase is ambiguous and does not clarify the role of the optional database parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, succinct and front-loaded with the core action. It contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the output schema exists and the tool is simple, the description leaves ambiguity about the database scope and parameter usage. It is marginally adequate but lacks important context for an agent to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Parameter documentation coverage is 0%, and the description does not explain the `table` and `database` parameters. It fails to mention required/optional status, formats, or defaults, which the schema alone must convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Return column metadata for a table', clearly identifying a specific verb (return) and resource (column metadata). The scope 'in one configured database' distinguishes this from sibling tools that list databases or perform data operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance about when to use this tool versus alternatives. It does not mention prerequisites, use cases, or when not to use it, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

executeC

Execute one guarded DML statement when writes are explicitly enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes
paramsNo
databaseNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It mentions 'guarded' and 'when writes are explicitly enabled', hinting at safety mechanisms, but does not explain what 'guarded' entails, whether changes are reversible, permission requirements, or likely side effects. For a mutating tool, this leaves significant ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is one sentence with no filler. It conveys the core purpose in a compact, front-loaded manner. Every word earns its place, even if the content is thin.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters, no annotations, and a write side-effect, the description is too sparse. It does not explain the return format, error behavior, or how to use params/database. The 'guarded' and 'enabled' context helps but is insufficient for safe invocation of a DML tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description provides zero information about the parameters (sql, params, database). The agent cannot learn what 'params' means (e.g., bind parameters) or how 'database' is used, leaving the tool's invocation underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool executes a DML statement, which distinguishes it from the sibling 'query' tool (reads). The qualifiers 'guarded' and 'when writes are explicitly enabled' add context, though 'guarded' is somewhat vague. It identifies the action and resource sufficiently.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is for write operations ('DML statement') and is conditioned on writes being enabled, but it does not explicitly contrast with 'query' or state when NOT to use it. The context is clear enough for a user already familiar with the tool set, but lacks explicit alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exportA

Stream a SELECT to Parquet and register it as a Sandbox Dataset.

The destination session and acting identity come from process-level environment variables, never from model-supplied tool arguments.

ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes
paramsNo
databaseNo
dataset_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It discloses a critical behavioral trait: 'The destination session and acting identity come from process-level environment variables, never from model-supplied tool arguments.' This goes beyond the generic action and warns against passing identity arguments. However, it does not mention potential side effects like overwriting existing datasets.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no redundant information. The first states the action, the second adds an essential behavioral note. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficient for a simple tool but not for one with 4 undocumented parameters and no annotations. It fails to explain optional parameters, database usage, or any prerequisites. The behavioral note is valuable but does not compensate for missing parameter guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no explanatory detail for any of the four parameters (sql, params, database, dataset_name). It only hints that sql is a SELECT query, but params and database are entirely unexplained. This is a major gap for tool invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: 'Stream a SELECT to Parquet and register it as a Sandbox Dataset.' This distinguishes it from sibling tools like query/execute (which run queries) and list_databases/describe_table (which inspect metadata).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool's use case: exporting query results to a persistent dataset. It clearly differentiates from siblings by its purpose, but does not explicitly state when not to use it or list alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_databasesA

List configured database ids without exposing DSNs or credentials.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It adds a meaningful security guarantee that DSNs and credentials are not exposed, and 'list' implies a read-only action. However, it does not explicitly discuss side effects or authentication requirements, though these are less critical for a zero-parameter read-only listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the action and resource. Every word earns its place, and there is no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, zero parameters, and the presence of an output schema, the description is reasonably complete. It states the scope and a key security constraint, though it could have added a hint about use cases or ordering, but these are not essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema is empty and the baseline is 4. The description clarifies that the output is limited to database ids, which adds semantic value beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists configured database ids, using a specific verb and resource. It also distinguishes itself from siblings like describe_table, query, execute, and export, which target tables or data operations rather than database enumeration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used to enumerate available database ids, which is a natural first step before using other tools, but it does not explicitly state when to use this tool versus alternatives. No exclusions or alternative references are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queryA

Run a bounded read-only SELECT. Use export for bulk pulls.

ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes
limitNo
paramsNo
databaseNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It discloses key safety traits: read-only (no mutation) and bounded (limited results). This is valuable context, though it doesn't specify default limits or error behavior, which is acceptable for a simple query tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loaded with the core action. 'Bounded' and 'read-only' add essential context, and the export alternative is stated efficiently. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and the tool is straightforward, the description is nearly sufficient. However, it doesn't clarify how 'bounded' maps to the limit parameter or what the database parameter does, leaving some ambiguity. It provides a basic but incomplete picture for a 4-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description mentions no parameters. It does not explain the sql, limit, params, or database parameters, leaving the agent without semantic guidance beyond type names. This is a major gap for a tool with multiple parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool runs a bounded read-only SELECT, clearly indicating a query operation. It distinguishes itself from siblings by specifying read-only (not execute) and bounded (not export for bulk pulls), making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use export for bulk pulls,' providing a concrete alternative for a specific scenario. The 'read-only' qualifier implies it's not for writes, giving clear when/when-not guidance relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.2.0
    • First observeddescribe_table
    • First observedexecute
    • First observedexport
    • First observedlist_databases
    • First observedquery

TDQS

A3.6/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: listing databases, describing table metadata, running bounded reads, executing guarded writes, and exporting bulk data. Query and export both run SELECTs, but the descriptions clearly differentiate interactive reads from bulk dataset creation.

Naming Consistency3/5

Naming is mixed: list_databases and describe_table follow a verb_noun pattern with underscores, while query, execute, and export are single verbs with no underscore. The convention is not consistent across the set, though all names are simple and readable.

Tool Count5/5

Five tools is well-scoped for a database MCP server, covering exploration, schema inspection, querying, writing, and bulk export. Each tool earns its place without redundancy or omission.

Completeness4/5

The surface covers the core database lifecycle: list databases, describe tables, query, write, and export. A minor gap is the lack of a list_tables tool, but agents can work around it by querying information_schema or using describe_table directly.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables direct interaction with MySQL databases through SQL queries, table exploration, and schema inspection with support for prepared statements and connection pooling.
    -
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to safely query MySQL databases with read-only access, featuring SQL injection protection, connection pooling, and automatic query limits for secure database exploration.
    4
    72 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables Claude Desktop to interact with MySQL databases through secure query execution, schema discovery, and multi-database support with configurable read/write permissions and built-in SQL injection protection.
    158 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to securely connect to and manage MySQL databases with support for multiple database connections, complete CRUD operations, schema inspection, and dynamic connection management through natural language.
    4 npm
    81
    MIT