StarRocks MCP Server
OfficialStarRocks 공식 MCP 서버
StarRocks MCP 서버는 AI 어시스턴트와 StarRocks 데이터베이스를 연결하는 다리 역할을 합니다. 복잡한 클라이언트 측 설정 없이도 SQL 직접 실행, 데이터베이스 탐색, 차트를 통한 데이터 시각화, 그리고 상세한 스키마/데이터 개요 검색이 가능합니다.
특징
직접 SQL 실행:
SELECT쿼리(read_query) 및 DDL/DML 명령(write_query)을 실행합니다.데이터베이스 탐색: 데이터베이스와 테이블을 나열하고, 테이블 스키마를 검색합니다(
starrocks://resources).시스템 정보:
proc://리소스 경로를 통해 내부 StarRocks 메트릭과 상태에 액세스합니다.자세한 개요: 열 정의, 행 수, 샘플 데이터를 포함하여 테이블(
table_overview)이나 전체 데이터베이스(db_overview)에 대한 포괄적인 요약을 얻습니다.데이터 시각화: 쿼리를 실행하고 결과에서 바로 Plotly 차트를 생성합니다(
query_and_plotly_chart).지능형 캐싱: 테이블 및 데이터베이스 개요를 메모리에 캐시하여 반복 요청 속도를 높입니다. 필요 시 캐시를 우회할 수 있습니다.
유연한 구성: 환경 변수를 통해 연결 세부 정보와 동작을 설정합니다.
Related MCP server: starrocks-mcp
구성
MCP 서버는 일반적으로 MCP 호스트를 통해 실행됩니다. StarRocks MCP 서버 프로세스를 시작하는 방법을 지정하는 구성이 호스트로 전달됩니다.
설치된 패키지에서 uv 사용:
지엑스피1
로컬 디렉토리에서 uv 사용(개발용):
{
"mcpServers": {
"mcp-server-starrocks": {
"command": "uv",
"args": [
"--directory",
"path/to/mcp-server-starrocks", // <-- Update this path
"run",
"mcp-server-starrocks"
],
"env": {
"STARROCKS_HOST": "default localhost",
"STARROCKS_PORT": "default 9030",
"STARROCKS_USER": "default root",
"STARROCKS_PASSWORD": "default empty",
"STARROCKS_DB": "default empty",
"STARROCKS_OVERVIEW_LIMIT": "default 20000"
}
}
}
}환경 변수:
STARROCKS_HOST: (선택 사항) StarRocks FE 서비스의 호스트 이름 또는 IP 주소입니다. 기본값은localhost입니다.STARROCKS_PORT: (선택 사항) StarRocks FE 서비스의 MySQL 프로토콜 포트입니다. 기본값은9030입니다.STARROCKS_USER: (선택 사항) StarRocks 사용자 이름입니다. 기본값은root입니다.STARROCKS_PASSWORD: (선택 사항) StarRocks 비밀번호입니다. 기본값은 빈 문자열입니다.STARROCKS_DB: (선택 사항) 도구 인수 또는 리소스 URI에 지정되지 않은 경우 사용할 기본 데이터베이스입니다. 설정된 경우 연결 시 이 데이터베이스를USE하려고 시도합니다.table_overview및db_overview와 같은 도구는 인수에 데이터베이스 부분이 생략된 경우 이 데이터베이스를 사용합니다. 기본값은 비어 있음(기본 데이터베이스 없음)입니다.STARROCKS_OVERVIEW_LIMIT: (선택 사항) 개요 도구(table_overview,db_overview)에서 캐시를 채우기 위해 데이터를 가져올 때 생성되는 총 텍스트의 대략적인 문자 수 제한입니다. 이 설정은 매우 큰 스키마나 여러 테이블에서 과도한 메모리 사용을 방지하는 데 도움이 됩니다. 기본값은20000입니다.
구성 요소
도구
read_query설명: ResultSet을 반환하는 SELECT 쿼리나 다른 명령(예:
SHOW,DESCRIBE)을 실행합니다.입력:
{ "query": "SQL query string" }출력: 헤더 행과 행 개수 요약을 포함한 CSV 형식의 쿼리 결과를 담은 텍스트 콘텐츠입니다. 실패 시 오류 메시지를 반환합니다.
write_query설명: ResultSet을 반환하지 않는 DDL(
CREATE,ALTER,DROP), DML(INSERT,UPDATE,DELETE) 또는 기타 StarRocks 명령을 실행합니다.입력:
{ "query": "SQL command string" }출력: 성공을 확인하는 텍스트 콘텐츠(예: "쿼리 성공, X개 행 영향 받음") 또는 오류를 보고합니다. 성공 시 변경 사항이 자동으로 커밋됩니다.
query_and_plotly_chart설명: SQL 쿼리를 실행하고, 결과를 Pandas DataFrame에 로드하고, 제공된 Python 표현식을 사용하여 Plotly 차트를 생성합니다. 지원되는 UI에서 시각화하도록 설계되었습니다.
입력:
{ "query": "SQL query to fetch data", "plotly_expr": "Python expression string using 'px' (Plotly Express) and 'df' (DataFrame). Example: 'px.scatter(df, x=\"col1\", y=\"col2\")'" }출력: 다음이 포함된 목록:
TextContent: DataFrame의 텍스트 표현과 차트가 UI 표시용이라는 메모입니다.ImageContent: 생성된 Plotly 차트를 base64 PNG 이미지(image/png)로 인코딩합니다. 실패하거나 쿼리에서 데이터가 반환되지 않으면 텍스트 오류 메시지를 반환합니다.
table_overview설명: 특정 테이블의 개요를 가져옵니다(열(
DESCRIBE에서 가져온 값), 총 행 수, 샘플 행(LIMIT 3)).refresh가 true가 아니면 메모리 내 캐시를 사용합니다.입력:
{ "table": "Table name, optionally prefixed with database name (e.g., 'db_name.table_name' or 'table_name'). If database is omitted, uses STARROCKS_DB environment variable if set.", "refresh": false // Optional, boolean. Set to true to bypass the cache. Defaults to false. }출력: 서식이 적용된 개요(열, 행 개수, 샘플 데이터) 또는 오류 메시지를 포함하는 텍스트 콘텐츠입니다. 캐시된 결과에는 이전 오류가 포함된 경우 포함됩니다.
db_overview설명: 지정된 데이터베이스 내 모든 테이블에 대한 개요(열, 행 수, 샘플 행)를 가져옵니다.
refresh참이 아닌 경우 각 테이블에 대해 테이블 수준 캐시를 사용합니다.입력:
{ "db": "database_name", // Optional if STARROCKS_DB env var is set. "refresh": false // Optional, boolean. Set to true to bypass the cache for all tables in the DB. Defaults to false. }출력: 데이터베이스에서 발견된 모든 테이블에 대한 연결 개요를 헤더로 구분하여 포함하는 텍스트 콘텐츠입니다. 데이터베이스에 액세스할 수 없거나 데이터베이스 테이블에 테이블이 없는 경우 오류 메시지를 반환합니다.
자원
직접 자원
starrocks:///databases설명: 구성된 사용자가 액세스할 수 있는 모든 데이터베이스를 나열합니다.
동등한 쿼리:
SHOW DATABASESMIME 유형:
text/plain
리소스 템플릿
starrocks:///{db}/{table}/schema설명: 특정 테이블의 스키마 정의를 가져옵니다.
동등한 쿼리:
SHOW CREATE TABLE {db}.{table}MIME 유형:
text/plain
starrocks:///{db}/tables설명: 특정 데이터베이스 내의 모든 테이블을 나열합니다.
동등한 쿼리:
SHOW TABLES FROM {db}MIME 유형:
text/plain
proc:///{+path}설명: Linux의
/proc명령과 유사하게 StarRocks 내부 시스템 정보에 접근합니다.path매개변수는 원하는 정보 노드를 지정합니다.동등한 쿼리:
SHOW PROC '/{path}'MIME 유형:
text/plain일반적인 경로:
/frontends- FE 노드에 대한 정보./backends- BE 노드에 대한 정보(클라우드 네이티브가 아닌 배포의 경우)./compute_nodes- CN 노드에 대한 정보(클라우드 네이티브 배포용)./dbs- 데이터베이스에 대한 정보./dbs/<DB_ID>- ID별 특정 데이터베이스에 대한 정보./dbs/<DB_ID>/<TABLE_ID>- ID별 특정 테이블에 대한 정보./dbs/<DB_ID>/<TABLE_ID>/partitions- 테이블에 대한 파티션 정보./transactions- 데이터베이스별로 그룹화된 거래 정보입니다./transactions/<DB_ID>- 특정 데이터베이스 ID에 대한 트랜잭션 정보./transactions/<DB_ID>/running- 데이터베이스 ID에 대한 트랜잭션 실행./transactions/<DB_ID>/finished- 데이터베이스 ID에 대한 완료된 트랜잭션./jobs- 비동기 작업(스키마 변경, 롤업 등)에 대한 정보입니다./statistic- 각 데이터베이스에 대한 통계./tasks- 에이전트 작업에 대한 정보./cluster_balance- 부하 분산 상태 정보./routine_loads- 루틴 로드 작업에 대한 정보입니다./colocation_group- Colocation Join 그룹에 대한 정보입니다./catalog- 구성된 카탈로그(예: Hive, Iceberg)에 대한 정보입니다.
프롬프트
이 서버에서는 아무것도 정의되지 않았습니다.
캐싱 동작
table_overview및db_overview도구는 메모리 내 캐시를 활용하여 생성된 개요 텍스트를 저장합니다.캐시 키는
(database_name, table_name)의 튜플입니다.table_overview호출되면 먼저 캐시를 확인합니다. 결과가 존재하고refresh매개변수가false(기본값)이면 캐시된 결과가 즉시 반환됩니다. 그렇지 않으면 StarRocks에서 데이터를 가져와 캐시에 저장한 다음 반환합니다.db_overview호출되면 데이터베이스의 모든 테이블을 나열한 다음,table_overview와 동일한 캐싱 로직(먼저 캐시를 확인하고, 필요한 경우 가져오며,refresh가false이거나 캐시 미스)을 사용하여 각 테이블 의 개요를 검색합니다.db_overview의refresh가true이면 해당 데이터베이스의 모든 테이블을 강제로 새로 고칩니다.STARROCKS_OVERVIEW_LIMIT환경 변수는 캐시를 채울 때 테이블당 생성되는 개요 문자열의 최대 길이에 대한 소프트 목표를 제공하여 메모리 사용량을 관리하는 데 도움이 됩니다.원래 페치 중에 발생한 오류 메시지를 포함한 캐시된 결과는 저장되어 후속 캐시 히트 시 반환됩니다.
데모

Available Tools
8 toolsanalyze_queryB
Analyze a query and get analyze result using query profile. Use set_session_db to set a per-session default database
| Name | Required | Description | Default |
|---|---|---|---|
| db | No | database | |
| sql | No | Query SQL | |
| uuid | No | Query ID, a string composed of 32 hexadecimal digits formatted as 8-4-4-4-12 |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions 'using query profile' but does not disclose whether the tool is read-only, requires authentication, has side effects, or what state (e.g., query must be previously executed) is needed. The behavioral traits are minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the purpose, and includes a concise usage hint. Every sentence adds value without unnecessary words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and the moderate complexity (3 parameters), the description covers the basic purpose and provides a hint about the database parameter. However, it does not clarify the difference between analyzing by SQL vs. UUID, or that the query may need to have been executed first. It is adequate but leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing baseline parameter descriptions. The description adds value by explaining that 'set_session_db' can set a per-session default database, indirectly clarifying that the 'db' parameter may be omitted if a default is set. This goes beyond the schema's simple 'database' label.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action ('analyze a query') and the resource ('query'), and mentions using 'query profile', which indicates the tool's specific function. However, the phrasing 'get analyze result' is slightly redundant, and it doesn't clearly distinguish from sibling tools like 'db_summary' or 'read_query'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes a hint to use 'set_session_db' for setting a default database, but provides no guidance on when to use this tool versus alternatives (e.g., 'read_query' or 'query_and_plotly_chart'). There is no mention of prerequisites, exclusions, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collect_query_dump_and_profileB
Run a query to get it's query dump and profile, output very large, need special tools to do further processing
| Name | Required | Description | Default |
|---|---|---|---|
| db | No | database | |
| query | Yes | query to execute |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the output is very large and needs special tools, which is useful. However, it does not mention other behavioral traits like destructiveness, permissions, or side effects, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the action. It is concise and to the point, though a bit more structure could improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 2 parameters, no output schema, and no nested objects, the description provides adequate context but does not explain what 'query dump' and 'profile' entail or the return format. It is minimally complete for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both parameters have descriptions: 'database' and 'query to execute'). The description adds no additional meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Run a query to get it's query dump and profile', which clearly identifies the verb (run) and resource (query dump and profile). It also mentions the output is very large, adding context. However, it does not differentiate from sibling tools like 'query_and_plotly_chart' or 'read_query'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by stating 'need special tools to do further processing', suggesting this tool is for large outputs requiring post-processing. However, it does not explicitly state when to use this tool versus alternatives, nor provide conditions to avoid.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
db_summaryA
Quickly get summary of a database with tables' schema and size information. Use set_session_db to set a per-session default database
| Name | Required | Description | Default |
|---|---|---|---|
| db | No | Database name. Optional: uses current database by default. | |
| limit | No | Output length limit in characters. Defaults to 10000. Higher values show more tables and details. | |
| refresh | No | Set to true to force refresh, ignoring cache. Defaults to false. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided; the description does not disclose if the tool is read-only, cached, or has side effects. It mentions the refresh parameter but does not explain caching behavior in text, leaving agents without key safety cues.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with an additional useful hint. It is front-loaded and contains no filler, efficiently conveying purpose and context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters and no annotations, the description covers the main use case but falls short on behavioral transparency. The presence of an output schema reduces the need to describe return values. Overall adequate but with gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for all parameters. The description adds some value by linking the 'db' parameter to set_session_db, but does not significantly expand on parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves a database summary including table schema and size, with a specific verb and resource. It distinguishes from siblings by mentioning set_session_db for default database context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates when to use (quickly get summary) and references set_session_db for setting a default database. However, it does not explicitly state when not to use or compare to siblings like read_query or table_overview.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_and_plotly_chartB
using sql query to extract data from database, then using python plotly_expr to generate a chart for UI to display. Use set_session_db to set a per-session default database
| Name | Required | Description | Default |
|---|---|---|---|
| db | No | database | |
| query | Yes | SQL query to execute | |
| format | No | chart output format, json|png|jpeg | jpeg |
| plotly_expr | Yes | a one function call expression, with 2 vars binded: `px` as `import plotly.express as px`, and `df` as dataframe generated by query `plotly_expr` example: `px.scatter(df, x="sepal_width", y="sepal_length", color="species", marginal_y="violin", marginal_x="box", trendline="ols", template="simple_white")` |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It explains the two-step process (query then chart) but does not mention side effects, errors, rate limits, or output format details. The description is simple but lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two sentences, the first clearly stating the main function. The second sentence provides a related tip but is somewhat tangential. It is well-structured and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description is fairly complete for a combined query-chart tool. It explains the process and mentions a prerequisite. However, it lacks details on output format or error handling, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions, so baseline is 3. The description adds context about the overall workflow but does not elaborate on individual parameters beyond the schema. The mention of set_session_db is peripheral.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it uses SQL query to extract data and then generates a chart with plotly_expr. It distinguishes from siblings like read_query (which only returns data) by explicitly mentioning chart generation. However, it could be more precise by contrasting with other query tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions using set_session_db to set a default database, which is helpful but does not provide guidance on when to use this tool over its siblings (e.g., read_query for data only, analyze_query for analysis). No explicit exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_queryA
Execute a SELECT query or commands that return a ResultSet. Set output_file to write the full result to disk instead of returning it inline (useful for large results).. Use set_session_db to set a per-session default database
| Name | Required | Description | Default |
|---|---|---|---|
| db | No | database | |
| query | Yes | SQL query to execute | |
| output_file | No | If set, write the full result to this file and return only a summary + small preview inline. Relative paths resolve against STARROCKS_MCP_OUTPUT_DIR (default: ~/.mcp-server-starrocks/output/). Absolute paths (and ~) are used as-is. Format is inferred from the file extension (.csv, .tsv, .json, .jsonl, .ndjson) unless output_format is given. NOTE: the file is written on the server's filesystem, which may not be the client machine in remote/http deployments. | |
| output_format | No | Override file format: csv|tsv|json|jsonl. If omitted, inferred from output_file extension; defaults to csv. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must fully disclose behavioral traits. It explains the output_file feature and notes that files are written on the server's filesystem, which is important for remote deployments. However, it does not explicitly state that the tool is read-only or discuss error handling or authentication.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with the purpose stated first. It includes two sentences plus a minor note, and every part adds value. The only flaw is an extra period after 'inline'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks details about the return structure when output_file is not used. It mentions returning 'inline' but does not specify the format or content (e.g., rows, columns). Given there is no output schema, this information is crucial for correct usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds significant value beyond the schema for output_file and output_format, explaining path resolution, environment variables, and format inference. For db, it adds no extra meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Execute a SELECT query or commands that return a ResultSet', which provides a specific verb and resource. It distinguishes this tool from siblings like write_query and analyze_query by focusing on read-only queries that produce a result set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives (e.g., write_query for modifications, analyze_query for explaining). The only instruction is to use set_session_db for default database, which is a side note, not a usage guideline for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_session_dbA
Set or clear the default database for THIS MCP session. Subsequent tool calls without an explicit db argument will use this database. Pass an empty string or null to clear and fall back to the server's global default. Returns the new effective default for this session.
| Name | Required | Description | Default |
|---|---|---|---|
| db | No | Database name to set as the per-session default. Empty/null clears the override. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the burden. It discloses that setting affects subsequent calls without explicit db argument, clarifies clearing behavior, and states the return value. No behavioral contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, no redundant words. Every sentence adds critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the single parameter and no annotations, the description is fully complete. It explains purpose, usage, parameter semantics, and return value. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds value by explaining that empty string or null clears the override, which is not explicitly in the schema. It clarifies the parameter's effect beyond the bare description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool sets or clears the default database for the MCP session, using specific verbs ('set', 'clear', 'fall back'). It distinguishes from sibling query tools by focusing on session state management.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use (set default) and how to clear (empty/null). While it doesn't explicitly state when not to use or list alternatives, the context of sibling tools makes the usage clear. The guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
table_overviewA
Get an overview of a specific table: columns, sample rows (up to 3), and total row count. Uses cache unless refresh=true. Use set_session_db to set a per-session default database
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | Table name, optionally prefixed with database name (e.g., 'db_name.table_name'). If database is omitted, uses the default database. | |
| refresh | No | Set to true to force refresh, ignoring cache. Defaults to false. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses caching behavior and refresh option, which is important for understanding tool behavior. No destructive actions are implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no fluff. First sentence clearly states the tool's action and output, second provides important context about caching and database setup. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity (2 params, output schema exists), the description covers key aspects: output components, caching, default database. Minor omission like error behavior or limit on sample rows is compensated by output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description does not add semantic meaning beyond the schema descriptions. The parameter details are fully covered by the schema, so description adds no extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it provides an overview of a table including columns, sample rows, and row count. However, it does not explicitly differentiate from sibling tools like read_query or analyze_query, which are for querying rather than generating a summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context on caching and default database setup via set_session_db, but lacks explicit guidance on when to use this tool versus alternatives (e.g., read_query for raw data).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_queryA
Execute a DDL/DML or other StarRocks command that do not have a ResultSet. Use set_session_db to set a per-session default database
| Name | Required | Description | Default |
|---|---|---|---|
| db | No | database | |
| query | Yes | SQL to execute |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description covers non-ResultSet nature but lacks details on side effects, permissions, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with core purpose, no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Missing output schema; description doesn't specify return format or error handling, but purpose and parameters are adequately covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers both parameters; description adds value by suggesting set_session_db for default database, enhancing understanding of db parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it executes DDL/DML commands without ResultSet, distinguishing it from sibling tools like read_query which returns results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions using set_session_db for default database but does not explicitly state when to use vs alternatives or when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.4.0- Changed
read_query2 fields changed- added
Input schema / properties / output_fileAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "If set, write the full result to this file and return only a summary + small preview inline. Relative paths resolve against STARROCKS_MCP_OUTPUT_DIR (default: ~/.mcp-server-starrocks/output/). Absolute paths (and ~) are used as-is. Format is inferred from the file extension (.csv, .tsv, .json, .jsonl, .ndjson) unless output_format is given. NOTE: the file is written on the server's filesystem, which may not be the client machine in remote/http deployments." +} - added
Input schema / properties / output_formatAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Override file format: csv|tsv|json|jsonl. If omitted, inferred from output_file extension; defaults to csv." +}
- Added
set_session_db
8 tool updates
v0.3.0- Changed
analyze_query12 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / dbAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "database" +} - added
Input schema / properties / sql / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / sql / defaultAdded value: +null - removed
Input schema / properties / sql / titleRemoved value: -"Sql" - removed
Input schema / properties / sql / typeRemoved value: -"string" - added
Input schema / properties / uuid / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / uuid / defaultAdded value: +null - removed
Input schema / properties / uuid / titleRemoved value: -"Uuid" - removed
Input schema / properties / uuid / typeRemoved value: -"string" - removed
Input schema / requiredRemoved value: -[ - "uuid", - "sql" -] - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "type": "string" + } + }, + "required": [ + "result" + ], + "type": "object", + "x-fastmcp-wrap-result": true +}
- Added
collect_query_dump_and_profile - Removed
db_overview - Added
db_summary - Changed
query_and_plotly_chart5 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / dbAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "database" +} - added
Input schema / properties / formatAdded value: +{ + "default": "jpeg", + "description": "chart output format, json|png|jpeg", + "type": "string" +} - removed
Input schema / properties / plotly_expr / titleRemoved value: -"Plotly Expr" - removed
Input schema / properties / query / titleRemoved value: -"Query"
- Changed
read_query3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / dbAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "database" +} - removed
Input schema / properties / query / titleRemoved value: -"Query"
- Changed
table_overview4 fields changed- added
Input schema / additionalPropertiesAdded value: +false - removed
Input schema / properties / refresh / titleRemoved value: -"Refresh" - removed
Input schema / properties / table / titleRemoved value: -"Table" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "type": "string" + } + }, + "required": [ + "result" + ], + "type": "object", + "x-fastmcp-wrap-result": true +}
- Changed
write_query3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / dbAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "database" +} - removed
Input schema / properties / query / titleRemoved value: -"Query"
6 tool updates
v1.0.0- First observed
analyze_query - First observed
db_overview - First observed
query_and_plotly_chart - First observed
read_query - First observed
table_overview - First observed
write_query
TDQS
Scored across 8 tools
Each tool targets a distinct function: query execution, analysis, chart generation, schema summary, etc. There is slight overlap between read_query and query_and_plotly_chart, but their outputs differ (raw data vs chart), and descriptions clarify the distinction. No major confusion.
Tool names follow no consistent pattern: some are verb_noun (e.g., read_query), some are noun_verb (e.g., db_summary), some include conjunctions (query_and_plotly_chart), and verbs vary (analyze, collect, set, write). This inconsistency may confuse an agent trying to infer tool purposes from naming.
With 8 tools, the server is well-scoped for a database MCP server covering querying, analysis, schema browsing, and charting. Each tool serves a clear purpose without being excessive or minimal.
The tool set covers core database interactions: query (read and write), analysis, schema overview, and charting. A minor gap is the absence of a tool to list all databases, but db_summary and set_session_db partially address this. Overall, the surface is reasonably complete for the stated purpose.
Maintenance
Related MCP Connectors
Ask data questions in natural language. Get SQL, insights, and charts from your databases.
Safe, read-only Postgres and MySQL access for AI agents. Audit log + column-level controls.
Draxlr's remote MCP server connects AI assistants to your SQL databases and dashboards. Explore schemas, run read-only queries, manage saved queries and dashboards, and export results, all with row-level security so each user sees only their own data.
Query 40 databases from Claude, ChatGPT, or Cursor — on any device. Read-only, encrypted, audited.
Related MCP Servers
- FlicenseBqualityDmaintenanceAn implementation of the Model Context Protocol that provides AI clients with intelligent diagnosis and analysis capabilities for StarRocks databases. It enables users to execute SQL queries, monitor storage health, and analyze performance issues through natural language interfaces.12-
- AlicenseAqualityDmaintenanceA read-only MCP server that enables users to query and explore StarRocks databases through AI assistants like Claude. It supports SQL execution, schema discovery, and secure LDAP authentication for data analysis and metadata exploration.41MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI assistants to query databases using natural language, with automatic schema discovery and SQL compilation.451 npm3,173Apache 2.0
- AlicenseAqualityDmaintenanceEnables AI assistants to connect to and interact with PostgreSQL, MySQL, SQLite, and MongoDB databases through natural language, supporting schema exploration, query execution, data export, and more.13MIT