biz-db-mcp
This server provides controlled, read-first MySQL access over MCP: browse configured databases, explore schemas, run bounded read-only queries, perform guarded DML writes, and export data to Sandbox datasets.
List configured database ids without exposing DSNs or credentials (
list_databases).Inspect table schemas and list tables/views with pattern filters (
describe_table,list_tables).Run read-only
SELECT/WITHqueries with per-database row limits, timeouts, parameter binding,offsetpagination, anddry_runpreflight (query).Get
EXPLAINplans without executing the query (explain).Execute a single guarded
INSERT/UPDATE/DELETE/REPLACEwhen writes are enabled, with WHERE requirements and max-affected-row protections, plusdry_runestimation (execute).Execute up to 100 DML statements in one transaction with optional atomic rollback (
execute_many).Stream a SELECT result to a Sandbox Dataset as Parquet/CSV/JSONL, with the destination identity taken from environment variables (
export).Support multiple databases with per-database overrides, SQL audit hashing, and optional DBPM password bridging.
Provides database operations for MySQL, including listing databases, describing tables, running read-only queries, executing protected single-statement DML writes, and exporting query results to Parquet for Sandbox Dataset uploads.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@biz-db-mcpshow me the schema for the orders table"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
biz-db-mcp
一个面向 Agent 的 MySQL MCP Server,支持多数据库、只读查询、受保护的 DML 写入,以及批量导出到 Sandbox Dataset。
v0.2 灵活性升级:按库覆盖配置、WHERE 白名单、
list_tables/explain勘探、dry_run预检、execute_many批量事务、offset分页与csv/jsonl导出。
设计边界
默认只读:
BIZDB_WRITE_ENABLED默认关闭,单个数据库还必须显式allow_writes: true。query与export只允许单条SELECT/WITH ... SELECT;execute只允许单条INSERT、UPDATE、DELETE或REPLACE,execute_many允许单事务批量(最多 100 条)。UPDATE/DELETE默认必须带WHERE,可通过BIZDB_WRITE_REQUIRE_WHERE_EXCEPT_TABLES按表名 glob 放宽(如temp_*,cache_*);写入在事务中执行,并受最大影响行数限制(支持按库覆盖)。export使用服务端游标分批写 Parquet/CSV/JSONL,再上传 Sandbox;身份从进程环境注入,模型不能指定目标 session。SQL 审计只记录 SHA-256 和长度,不记录原始 SQL、参数或密码。
所有限流与超时均支持按库覆盖,全局配置为 fallback。
Related MCP server: MySQL ReadOnly MCP Server
目录结构
biz_db_mcp/
config.py # 环境变量、多数据库、写入和 DBPM 配置(含按库覆盖与文件加载)
db.py # MySQL 连接、查询、写入、勘探(describe/list_tables/explain/批量)
dbpm.py # DBPM 外部桥接命令适配器
sql_guard.py # SELECT/DML 语句边界
export.py # 流式 Parquet/CSV/JSONL 导出
sandbox_client.py # Dataset 上传(多格式 MIME)
audit.py # 结构化审计日志
server.py # MCP 工具注册(8 个工具)
tests/ # 单元测试配置
单库兼容配置:BIZDB_MYSQL_DSN=mysql://user:password@host:3306/database。
多库使用 BIZDB_DATABASES_JSON,格式可以是映射或 connections 数组:
{"connections":[
{"id":"primary","dsn":"mysql://user:password@db-a/app","query_max_rows":200},
{"id":"reporting","dsn":"mysql://user:password@db-b/reporting","allow_writes":false,"query_max_rows":5000}
]}也支持文件加载:BIZDB_DATABASES_FILE=/etc/biz-db.json(内容与 BIZDB_DATABASES_JSON 同格式,优先级:BIZDB_DATABASES_JSON > BIZDB_DATABASES_FILE > BIZDB_MYSQL_DSN)。
BIZDB_DEFAULT_DATABASE 选择默认库。写入还需要 BIZDB_WRITE_ENABLED=true,并在目标库配置 allow_writes: true;可用 BIZDB_WRITE_MAX_AFFECTED_ROWS 和 BIZDB_WRITE_REQUIRE_WHERE 调整保护阈值。按库覆盖示例:
{"id":"writer","dsn":"mysql://...","allow_writes":true,"write_max_affected_rows":5000,"write_require_where":false,"query_timeout_seconds":30}支持的按库覆盖字段:query_max_rows、query_timeout_seconds、export_max_rows、export_timeout_seconds、write_max_affected_rows、write_require_where(null 表示继承全局)。
WHERE 白名单:BIZDB_WRITE_REQUIRE_WHERE_EXCEPT_TABLES=temp_*,cache_logs(逗号或空格分隔,glob 匹配,大小写不敏感),命中白名单的表允许 UPDATE/DELETE 不带 WHERE。
工具一览(8 个)
工具 | 说明 | 灵活性增强 |
| 列出已配置库,不暴露 DSN | 按库覆盖后仍脱敏返回 |
| 返回表结构,支持 | 支持 |
| 列出库内表/视图,支持 | 新增 |
| 受控只读 SELECT, | 新增 |
| 返回 | 新增 |
| 单条 DML 事务提交 | 新增 |
| 多条 DML 单事务批量(≤100 条) | 新增, |
| 流式导出到 Sandbox Dataset | 新增 |
execute 仍为单语句 DML 接口,不提供 DDL、事务控制或多语句拼接;批量场景请用 execute_many。
DBPM(可选)
UPspec 中的 DBPM 是密码管理系统,不是 MCP 请求认证。Python 项目不内置私有 Java SDK;启用 BIZDB_DBPM_ENABLED=true 后,必须配置 BIZDB_DBPM_COMMAND、BIZDB_DBPM_SERVERS、BIZDB_DBPM_APP_NAME、BIZDB_DBPM_APP_KEY 和 BIZDB_DBPM_APP_IP。
桥接命令从 stdin 读取 JSON(包含 database、username、servers 等字段),只在 stdout 输出密码,非零退出码表示失败。DBPM 私钥通过 DBPM_APP_KEY 环境变量传给桥接进程,不出现在命令行参数中。数据库条目可用 dbpm_database/dbpm_user 覆盖 DBPM 查询键。
开发与运行
python3 -m pip install -e '.[dev]'
python3 -m pytest
biz-db-mcpSandbox 上传仍需配置 BIZDB_SANDBOX_BASE_URL、BIZDB_SANDBOX_API_TOKEN;导出身份需配置 BIZDB_SESSION_ID、BIZDB_ORG_ID、BIZDB_USER_ID。生产环境应使用 DBPM,不要把真实密码提交到配置、日志或测试夹具中。
灵活性使用示例
# 分页查询
query("SELECT id, name FROM users WHERE active=1", limit=50, offset=100)
# 预检(不查库)
query("SELECT * FROM orders WHERE status=%s", ["paid"], dry_run=True)
# -> {"dry_run": True, "effective_limit": 200, "effective_timeout": 10}
explain("SELECT * FROM orders WHERE status=%s", ["paid"])
list_tables(pattern="order%", include_views=False)
describe_table("reporting.orders")
# 写入预估
execute("UPDATE users SET active=0 WHERE id=%s", [123], dry_run=True)
# -> {"dry_run": True, "estimated_plan": [...] }
# 批量事务
execute_many([
{"sql": "INSERT INTO users(name) VALUES (%s)", "params": ["alice"]},
{"sql": "UPDATE users SET active=1 WHERE id=%s", "params": [123]}
], atomic=True)
# 多格式导出
export("SELECT * FROM orders", dataset_name="orders_2024", format="csv")
export("SELECT * FROM orders", dataset_name="orders_2024", format="jsonl")环境变量速查(新增标记 ★)
变量 | 说明 | 默认 |
| 多库 JSON 文件路径 | - |
| 免 WHERE 白名单(glob) | - |
单库按库覆盖 ★ |
| 继承全局 |
Available Tools
5 toolsdescribe_tableB
Return column metadata for a table in one configured database.
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | ||
| database | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavioral traits. It only mentions the return of metadata and a vague scope ('one configured database'), without addressing read-only status, error handling, or permissions. The phrase is ambiguous and does not clarify the role of the optional database parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, succinct and front-loaded with the core action. It contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the output schema exists and the tool is simple, the description leaves ambiguity about the database scope and parameter usage. It is marginally adequate but lacks important context for an agent to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Parameter documentation coverage is 0%, and the description does not explain the `table` and `database` parameters. It fails to mention required/optional status, formats, or defaults, which the schema alone must convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Return column metadata for a table', clearly identifying a specific verb (return) and resource (column metadata). The scope 'in one configured database' distinguishes this from sibling tools that list databases or perform data operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance about when to use this tool versus alternatives. It does not mention prerequisites, use cases, or when not to use it, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
executeC
Execute one guarded DML statement when writes are explicitly enabled.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | ||
| params | No | ||
| database | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It mentions 'guarded' and 'when writes are explicitly enabled', hinting at safety mechanisms, but does not explain what 'guarded' entails, whether changes are reversible, permission requirements, or likely side effects. For a mutating tool, this leaves significant ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one sentence with no filler. It conveys the core purpose in a compact, front-loaded manner. Every word earns its place, even if the content is thin.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters, no annotations, and a write side-effect, the description is too sparse. It does not explain the return format, error behavior, or how to use params/database. The 'guarded' and 'enabled' context helps but is insufficient for safe invocation of a DML tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides zero information about the parameters (sql, params, database). The agent cannot learn what 'params' means (e.g., bind parameters) or how 'database' is used, leaving the tool's invocation underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool executes a DML statement, which distinguishes it from the sibling 'query' tool (reads). The qualifiers 'guarded' and 'when writes are explicitly enabled' add context, though 'guarded' is somewhat vague. It identifies the action and resource sufficiently.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is for write operations ('DML statement') and is conditioned on writes being enabled, but it does not explicitly contrast with 'query' or state when NOT to use it. The context is clear enough for a user already familiar with the tool set, but lacks explicit alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exportA
Stream a SELECT to Parquet and register it as a Sandbox Dataset.
The destination session and acting identity come from process-level environment variables, never from model-supplied tool arguments.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | ||
| params | No | ||
| database | No | ||
| dataset_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It discloses a critical behavioral trait: 'The destination session and acting identity come from process-level environment variables, never from model-supplied tool arguments.' This goes beyond the generic action and warns against passing identity arguments. However, it does not mention potential side effects like overwriting existing datasets.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundant information. The first states the action, the second adds an essential behavioral note. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is sufficient for a simple tool but not for one with 4 undocumented parameters and no annotations. It fails to explain optional parameters, database usage, or any prerequisites. The behavioral note is valuable but does not compensate for missing parameter guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description adds no explanatory detail for any of the four parameters (sql, params, database, dataset_name). It only hints that sql is a SELECT query, but params and database are entirely unexplained. This is a major gap for tool invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: 'Stream a SELECT to Parquet and register it as a Sandbox Dataset.' This distinguishes it from sibling tools like query/execute (which run queries) and list_databases/describe_table (which inspect metadata).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool's use case: exporting query results to a persistent dataset. It clearly differentiates from siblings by its purpose, but does not explicitly state when not to use it or list alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_databasesA
List configured database ids without exposing DSNs or credentials.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It adds a meaningful security guarantee that DSNs and credentials are not exposed, and 'list' implies a read-only action. However, it does not explicitly discuss side effects or authentication requirements, though these are less critical for a zero-parameter read-only listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the action and resource. Every word earns its place, and there is no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, zero parameters, and the presence of an output schema, the description is reasonably complete. It states the scope and a key security constraint, though it could have added a hint about use cases or ordering, but these are not essential.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is empty and the baseline is 4. The description clarifies that the output is limited to database ids, which adds semantic value beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists configured database ids, using a specific verb and resource. It also distinguishes itself from siblings like describe_table, query, execute, and export, which target tables or data operations rather than database enumeration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to enumerate available database ids, which is a natural first step before using other tools, but it does not explicitly state when to use this tool versus alternatives. No exclusions or alternative references are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queryA
Run a bounded read-only SELECT. Use export for bulk pulls.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | ||
| limit | No | ||
| params | No | ||
| database | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It discloses key safety traits: read-only (no mutation) and bounded (limited results). This is valuable context, though it doesn't specify default limits or error behavior, which is acceptable for a simple query tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loaded with the core action. 'Bounded' and 'read-only' add essential context, and the export alternative is stated efficiently. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the tool is straightforward, the description is nearly sufficient. However, it doesn't clarify how 'bounded' maps to the limit parameter or what the database parameter does, leaving some ambiguity. It provides a basic but incomplete picture for a 4-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions no parameters. It does not explain the sql, limit, params, or database parameters, leaving the agent without semantic guidance beyond type names. This is a major gap for a tool with multiple parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool runs a bounded read-only SELECT, clearly indicating a query operation. It distinguishes itself from siblings by specifying read-only (not execute) and bounded (not export for bulk pulls), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use export for bulk pulls,' providing a concrete alternative for a specific scenario. The 'read-only' qualifier implies it's not for writes, giving clear when/when-not guidance relative to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.2.0- First observed
describe_table - First observed
execute - First observed
export - First observed
list_databases - First observed
query
TDQS
Scored across 5 tools
Each tool has a clear, distinct purpose: listing databases, describing table metadata, running bounded reads, executing guarded writes, and exporting bulk data. Query and export both run SELECTs, but the descriptions clearly differentiate interactive reads from bulk dataset creation.
Naming is mixed: list_databases and describe_table follow a verb_noun pattern with underscores, while query, execute, and export are single verbs with no underscore. The convention is not consistent across the set, though all names are simple and readable.
Five tools is well-scoped for a database MCP server, covering exploration, schema inspection, querying, writing, and bulk export. Each tool earns its place without redundancy or omission.
The surface covers the core database lifecycle: list databases, describe tables, query, write, and export. A minor gap is the lack of a list_tables tool, but agents can work around it by querying information_schema or using describe_table directly.
Maintenance
Related MCP Connectors
Safe, read-only Postgres and MySQL access for AI agents. Audit log + column-level controls.
PostgreSQL, MySQL, OpenAPI/Swagger, and shared Agent Memory with scoped access.
- dataOAuthco.thinair
Read-only PostgreSQL, MySQL, SQL Server access via MCP — 24 dialect-aware hosted tools.
Query 40 databases from Claude, ChatGPT, or Cursor — on any device. Read-only, encrypted, audited.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceEnables direct interaction with MySQL databases through SQL queries, table exploration, and schema inspection with support for prepared statements and connection pooling.-
- AlicenseAqualityDmaintenanceEnables AI assistants to safely query MySQL databases with read-only access, featuring SQL injection protection, connection pooling, and automatic query limits for secure database exploration.472 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables Claude Desktop to interact with MySQL databases through secure query execution, schema discovery, and multi-database support with configurable read/write permissions and built-in SQL injection protection.158 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to securely connect to and manage MySQL databases with support for multiple database connections, complete CRUD operations, schema inspection, and dynamic connection management through natural language.4 npm81MIT