redshift-utils-mcp
Servidor MCP de Redshift Utils
Descripción general
Este proyecto implementa un servidor de Protocolo de Contexto de Modelo (MCP) diseñado específicamente para interactuar con bases de datos de Amazon Redshift.
Conecta los Modelos de Lenguaje Grandes (LLM) o los asistentes de IA (como los de Claude, Cursor o aplicaciones personalizadas) con el almacén de datos de Redshift, lo que permite un acceso y una interacción seguros y estandarizados con los datos. Esto permite a los usuarios consultar datos, comprender la estructura de la base de datos y realizar operaciones de monitorización y diagnóstico mediante lenguaje natural o indicaciones basadas en IA.
Este servidor es para desarrolladores, analistas de datos o equipos que buscan integrar las capacidades de LLM directamente con su entorno de datos de Amazon Redshift de manera estructurada y segura.
Related MCP server: Redshift MCP Server
Tabla de contenido
Características
✨ Conexión segura a Redshift (a través de API de datos): se conecta a su clúster de Amazon Redshift mediante la API de datos de AWS Redshift a través de Boto3, aprovechando AWS Secrets Manager para credenciales administradas de forma segura a través de variables de entorno.
🔍 Descubrimiento de esquemas: expone recursos MCP para enumerar esquemas y tablas dentro de un esquema específico.
📊 Metadatos y estadísticas: proporciona una herramienta (
handle_inspect_table) para recopilar metadatos de tablas detallados, estadísticas (como tamaño, recuento de filas, sesgo, obsolescencia de las estadísticas) y estado de mantenimiento.Ejecución de consultas de solo lectura: ofrece una herramienta MCP segura (
handle_execute_ad_hoc_query) para ejecutar consultas SELECT arbitrarias en la base de datos Redshift, lo que permite la recuperación de datos en función de solicitudes LLM.📈 Análisis del rendimiento de consultas: incluye una herramienta (
handle_diagnose_query_performance) para recuperar y analizar el plan de ejecución, las métricas y los datos históricos de un ID de consulta específico.🔍 Inspección de tabla: proporciona una herramienta (
handle_inspect_table) para realizar una inspección integral de una tabla, incluido el diseño, el almacenamiento, el estado y el uso.🩺 Comprobación del estado del clúster: ofrece una herramienta (
handle_check_cluster_health) para realizar una evaluación básica o completa del estado del clúster mediante varias consultas de diagnóstico.🔒 Diagnóstico de bloqueo: proporciona una herramienta (
handle_diagnose_locks) para identificar e informar sobre la contención de bloqueos actuales y las sesiones de bloqueo.Monitoreo de carga de trabajo: incluye una herramienta (
handle_monitor_workload) para analizar patrones de carga de trabajo del clúster durante una ventana de tiempo, cubriendo WLM, consultas principales y uso de recursos.📝 Recuperación de DDL: ofrece una herramienta (
handle_get_table_definition) para recuperar la salida deSHOW TABLE(DDL) para una tabla específica.🛡️ Saneamiento de entrada: utiliza consultas parametrizadas a través del cliente de API de datos Boto3 Redshift cuando corresponde para mitigar los riesgos de inyección de SQL.
🧩 Interfaz MCP estandarizada: se adhiere a la especificación del Protocolo de contexto de modelo para una integración perfecta con clientes compatibles (por ejemplo, Claude Desktop, Cursor IDE, aplicaciones personalizadas).
Prerrequisitos
Software:
Python 3.8+
uv(gestor de paquetes recomendado)Git (para clonar el repositorio)
Infraestructura y acceso:
Acceso a un clúster de Amazon Redshift.
Una cuenta de AWS con permisos para usar la API de datos de Redshift (
redshift-data:*) y acceder al secreto de Secrets Manager especificado (secretsmanager:GetSecretValue).Una cuenta de usuario de Redshift cuyas credenciales se almacenan en AWS Secrets Manager. Este usuario necesita los permisos necesarios en Redshift para realizar las acciones habilitadas por este servidor (p. ej.,
CONNECTcon la base de datos,SELECTen las tablas de destino,SELECTen las vistas del sistema relevantes comopg_class,pg_namespace,svv_all_schemas,svv_tablesy `svv_table_info`). Se recomienda encarecidamente usar un rol con el principio de mínimo privilegio. Consulte Consideraciones de seguridad .
Cartas credenciales:
Los datos de su conexión a Redshift se gestionan mediante AWS Secrets Manager, y el servidor se conecta mediante la API de datos de Redshift. Necesita:
El identificador del grupo Redshift.
El nombre de la base de datos dentro del clúster.
El ARN del secreto de AWS Secrets Manager que contiene las credenciales de la base de datos (nombre de usuario y contraseña).
La región de AWS donde residen el clúster y el secreto.
Opcionalmente, un nombre de perfil de AWS si no se utilizan credenciales/región predeterminadas.
Estos detalles se configurarán a través de variables de entorno como se detalla en la sección Configuración .
Configuración
Configurar variables de entorno: Este servidor requiere las siguientes variables de entorno para conectarse a su clúster de Redshift mediante la API de datos de AWS. Puede configurarlas directamente en su shell mediante un archivo de servicio systemd, un archivo de entorno de Docker o creando un archivo .env en el directorio raíz del proyecto (si utiliza una herramienta como uv o python-dotenv que admita la carga desde .env ).
Ejemplo usando la exportación de shell:
export REDSHIFT_CLUSTER_ID="your-cluster-id"
export REDSHIFT_DATABASE="your_database_name"
export REDSHIFT_SECRET_ARN="arn:aws:secretsmanager:us-east-1:123456789012:secret:your-redshift-secret-XXXXXX"
export AWS_REGION="us-east-1" # Or AWS_DEFAULT_REGION
# export AWS_PROFILE="your-aws-profile-name" # OptionalEjemplo de archivo .env (ver .env.example ):
# .env file for Redshift MCP Server configuration
# Ensure this file is NOT committed to version control if it contains secrets. Add it to .gitignore.
REDSHIFT_CLUSTER_ID="your-cluster-id"
REDSHIFT_DATABASE="your_database_name"
REDSHIFT_SECRET_ARN="arn:aws:secretsmanager:us-east-1:123456789012:secret:your-redshift-secret-XXXXXX"
AWS_REGION="us-east-1" # Or AWS_DEFAULT_REGION
# AWS_PROFILE="your-aws-profile-name" # OptionalTabla de variables requeridas:
Nombre de la variable | Requerido | Descripción | Valor de ejemplo |
| Sí | Su identificador de clúster Redshift. |
|
| Sí | El nombre de la base de datos a la que conectarse. |
|
| Sí | ARN de AWS Secrets Manager para credenciales de Redshift. |
|
| Sí | Región de AWS para API de datos y Secrets Manager. |
|
| No | Alternativa a |
|
| No | Nombre de perfil de AWS que se utilizará desde su archivo de credenciales (~/.aws/...). |
|
Nota: asegúrese de que las credenciales de AWS utilizadas por Boto3 (a través del entorno, perfil o rol de IAM) tengan permisos para acceder al REDSHIFT_SECRET_ARN especificado y usar la API de datos de Redshift ( redshift-data:* ).
Uso
Conexión con Claude Desktop / Consola Antrópica:
Agregue el siguiente bloque de configuración a su archivo mcp.json . Ajuste command , args , env y workingDirectory según su método de instalación y configuración.
{
"mcpServers": {
"redshift-utils-mcp": {
"command": "uvx",
"args": ["redshift_utils_mcp"],
"env": {
"REDSHIFT_CLUSTER_ID":"your-cluster-id",
"REDSHIFT_DATABASE":"your_database_name",
"REDSHIFT_SECRET_ARN":"arn:aws:secretsmanager:...",
"AWS_REGION": "us-east-1"
}
}
}Conexión con Cursor IDE:
Inicie el servidor MCP localmente siguiendo las instrucciones de la sección Uso / Inicio rápido .
En Cursor, abra la Paleta de comandos (Cmd/Ctrl + Shift + P).
Escriba "Conectarse al servidor MCP" o navegue a la configuración de MCP.
Agregar una nueva conexión al servidor.
Seleccione el tipo de transporte
stdio.Introduzca el comando y los argumentos necesarios para iniciar el servidor (
uvx run redshift_utils_mcp). Asegúrese de que todas las variables de entorno necesarias estén disponibles para el comando que se está ejecutando.El cursor debe detectar el servidor y sus herramientas/recursos disponibles.
Recursos MCP disponibles
Patrón de URI de recurso | Descripción | Ejemplo de URI |
| Recupera el contenido sin procesar de un archivo de script SQL del directorio |
|
| Enumera todos los esquemas definidos por el usuario accesibles en la base de datos conectada. |
|
| Recupera los detalles de configuración actuales de Gestión de carga de trabajo (WLM). |
|
| Enumera todas las tablas y vistas accesibles dentro del |
|
Reemplace {script_path} y {schema_name} con los valores reales al realizar solicitudes. La accesibilidad a los esquemas/tablas depende de los permisos otorgados al usuario de Redshift, configurados mediante REDSHIFT_SECRET_ARN .
Herramientas MCP disponibles
Nombre de la herramienta | Descripción | Parámetros clave (obligatorios*) | Ejemplo de invocación |
| Realiza una evaluación del estado del clúster Redshift utilizando un conjunto de scripts SQL de diagnóstico. |
|
|
| Identifica contención de bloqueo activa y sesiones de bloqueo en el clúster. |
|
|
| Analiza el rendimiento de ejecución de una consulta específica, incluido el plan, las métricas y los datos históricos. |
|
|
| Ejecuta una consulta SQL arbitraria proporcionada por el usuario a través de la API de datos de Redshift. Diseñado como una vía de escape. |
|
|
| Recupera la declaración DDL (lenguaje de definición de datos) ( |
|
|
| Recupera información detallada sobre una tabla Redshift específica, que abarca el diseño, el almacenamiento, el estado y el uso. |
|
|
| Analiza los patrones de carga de trabajo del clúster durante un período de tiempo específico utilizando varios scripts de diagnóstico. |
|
|
HACER
[ ] Mejorar las opciones de aviso
[ ] Agregar soporte para más métodos de credenciales
[ ] Agregar soporte para Redshift Serverless
Contribuyendo
¡Agradecemos sus contribuciones! Por favor, sigan estas pautas.
Encontrar/Reportar problemas: Consulta la página de problemas de GitHub para ver errores o solicitudes de funciones. Si lo necesitas, puedes abrir un nuevo problema.
La seguridad es fundamental al proporcionar acceso a bases de datos a través de un servidor MCP. Tenga en cuenta lo siguiente:
🔒 Administración de credenciales: Este servidor utiliza AWS Secrets Manager a través de la API de datos de Redshift, lo cual es más seguro que almacenar las credenciales directamente en variables de entorno o archivos de configuración. Asegúrese de que las credenciales de AWS que utiliza Boto3 (a través del entorno, perfil o rol de IAM) se administren de forma segura y tengan los permisos mínimos necesarios. Nunca envíe sus credenciales de AWS ni archivos .env que contengan secretos al control de versiones.
🛡️ Principio de privilegios mínimos: Configure el usuario de Redshift cuyas credenciales se encuentran en AWS Secrets Manager con los permisos mínimos necesarios para la funcionalidad prevista del servidor. Por ejemplo, si solo se necesita acceso de lectura, otorgue únicamente los privilegios CONNECT y SELECT en los esquemas/tablas necesarios, y SELECT en las vistas del sistema requeridas. Evite usar usuarios con privilegios elevados, como admin o el superusuario del clúster.
Para obtener orientación sobre la creación de usuarios restringidos de Redshift y la administración de permisos, consulte el sitio oficial ( https://docs.aws.amazon.com/redshift/latest/mgmt/security.html ).
Licencia
Este proyecto está licenciado bajo la Licencia MIT. Consulte el archivo (LICENCIA) para más detalles.
Referencias
Este proyecto se basa en gran medida en la especificación del Protocolo de Contexto de Modelo .
Desarrollado utilizando el SDK oficial de MCP proporcionado por Model Context Protocol .
Utiliza el AWS SDK para Python ( Boto3 ) para interactuar con la API de datos de Amazon Redshift .
Muchos de los scripts SQL de diagnóstico están adaptados del excelente repositorio awslabs/amazon-redshift-utils .
Available Tools
7 toolshandle_check_cluster_healthA
Performs a health assessment of the Redshift cluster.
Executes a series of diagnostic SQL scripts concurrently based on the
specified level ('basic' or 'full'). Aggregates raw results or errors
from each script into a dictionary.
Args:
ctx: The MCP context object.
level: Level of detail: 'basic' for operational status, 'full' for
comprehensive table design/maintenance checks. Defaults to 'basic'.
time_window_days: Lookback period in days for time-sensitive checks
(e.g., queue waits, commit waits). Defaults to 1.
Returns:
A dictionary where keys are script names and values are either the raw
list of dictionary results from the SQL query or an Exception object
if that specific script failed.
Raises:
DataApiError: If a critical error occurs during script execution that
prevents gathering results (e.g., config error). Individual
script errors are captured within the returned dictionary.
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | basic | |
| time_window_days | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does well by disclosing key behavioral traits: concurrent execution of scripts, aggregation of results into a dictionary, error handling approach (individual script errors captured in dictionary vs. critical errors raised as DataApiError), and the distinction between basic and full diagnostic levels.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, execution behavior, args, returns, raises) and front-loaded with the core purpose. While comprehensive, some sentences could be more concise, such as the detailed explanation of the return dictionary which is slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and 0% schema description coverage, the description provides substantial context including purpose, parameters, return format, and error handling. However, it doesn't mention authentication requirements, rate limits, or potential side effects on the cluster, which would be helpful given the diagnostic nature.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description fully compensates by providing detailed semantic explanations for both parameters: 'level' options ('basic' for operational status, 'full' for comprehensive checks) and 'time_window_days' purpose (lookback period for time-sensitive checks like queue waits). It also mentions default values and provides concrete examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'performs a health assessment of the Redshift cluster' with specific verbs ('executes diagnostic SQL scripts', 'aggregates results') and distinguishes it from siblings by focusing on comprehensive cluster health rather than specific issues like locks, query performance, or table inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use different levels ('basic' for operational status, 'full' for comprehensive checks) and mentions time-sensitive checks, but doesn't explicitly state when to choose this tool over sibling tools like handle_diagnose_query_performance or handle_monitor_workload for similar health-related tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_diagnose_locksA
Identifies active lock contention in the cluster.
Fetches all current lock information and then filters it based on the
optional target PID, target table name, and minimum wait time.
Formats the results into a list of contention details and a summary.
Args:
ctx: The MCP context object.
target_pid: Optional: Filter results to show locks held by or waited
for by this specific process ID (PID).
target_table_name: Optional: Filter results for locks specifically on
this table name (schema qualification recommended
if ambiguous).
min_wait_seconds: Minimum seconds a lock must be in a waiting state
to be included. Defaults to 5.
Returns:
A list of dictionaries, where each dictionary represents a row
from the lock contention query result.
Raises:
DataApiError: If fetching the initial lock information fails.
| Name | Required | Description | Default |
|---|---|---|---|
| min_wait_seconds | No | ||
| target_pid | No | ||
| target_table_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does well by explaining the tool's multi-step behavior: fetching all lock information, applying optional filters, formatting results into list+summary structure, and potential error conditions (DataApiError). It doesn't mention permissions, rate limits, or side effects, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and well-structured with clear sections (purpose, args, returns, raises). While efficient, the parameter explanations could be slightly more concise, and the purpose statement could be more front-loaded before diving into implementation details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a diagnostic tool with 3 parameters, no annotations, and no output schema, the description provides good coverage: clear purpose, parameter semantics, return format (list of dictionaries), and error conditions. It could improve by explaining the summary structure or providing example output, but overall it's reasonably complete given the context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by providing detailed semantic explanations for all three parameters: target_pid (filter by process ID), target_table_name (filter by table with schema qualification note), and min_wait_seconds (minimum waiting time with default). The descriptions add meaningful context beyond basic schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('identifies', 'fetches', 'filters', 'formats') and resource ('active lock contention in the cluster'). It distinguishes itself from siblings by focusing specifically on lock diagnostics rather than general health, performance, or table operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through parameter explanations (filtering by PID, table name, wait time) but doesn't explicitly state when to use this tool versus alternatives like handle_check_cluster_health or handle_diagnose_query_performance. No explicit when-not-to-use guidance or named alternatives are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_diagnose_query_performanceA
Analyzes a specific query's execution performance.
Fetches query text, execution plan, metrics, alerts, compilation info,
skew details, and optionally historical run data. Uses a formatting
utility to synthesize this into a structured report with potential issues
and recommendations.
Args:
ctx: The MCP context object.
query_id: The numeric ID of the Redshift query to analyze.
compare_historical: Fetch performance data for previous runs of the
same query text. Defaults to True.
Returns:
A dictionary conforming to DiagnoseQueryPerformanceResult structure:
- On success: Contains detailed performance breakdown, issues, recommendations.
- On query not found: Raises QueryNotFound exception.
- On other errors: Raises DataApiError or similar for FastMCP to handle.
Raises:
DataApiError: If a critical error occurs during script execution or parsing.
QueryNotFound: If the specified query_id cannot be found in key tables.
| Name | Required | Description | Default |
|---|---|---|---|
| compare_historical | No | ||
| query_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and does so well. It describes what data gets fetched, how it's synthesized into a structured report, and documents specific error conditions (QueryNotFound, DataApiError). It also mentions the formatting utility and the tool's ability to optionally fetch historical data. While it doesn't mention rate limits or authentication needs, it provides substantial behavioral context for a diagnostic tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and well-structured with clear sections: purpose statement, what it fetches, how it processes data, args documentation, returns documentation, and raises documentation. Every sentence earns its place, though the returns section could be slightly more concise. The information is front-loaded with the core purpose stated first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a diagnostic tool with 2 parameters, no annotations, and no output schema, the description provides substantial context. It explains what data gets collected, how it's processed, parameter meanings, and error conditions. The main gap is the lack of detail about the exact structure of the returned dictionary or what specific metrics/alerts are examined, but given the tool's complexity, this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by providing detailed parameter semantics. It explains that query_id is 'the numeric ID of the Redshift query to analyze' and that compare_historical controls whether to 'fetch performance data for previous runs of the same query text' with its default value. This adds crucial meaning beyond the bare schema types (integer, boolean).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('analyzes', 'fetches', 'synthesizes') and resources ('query's execution performance', 'query text, execution plan, metrics, alerts, compilation info, skew details, historical run data'). It distinguishes from sibling tools like handle_check_cluster_health or handle_diagnose_locks by focusing specifically on query performance analysis rather than cluster health or lock diagnosis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: when you need to analyze a specific query's performance with detailed metrics and recommendations. It doesn't explicitly state when NOT to use it or name specific alternatives among siblings, but the context is sufficiently clear for an agent to understand this is for query performance diagnosis rather than general cluster monitoring or table inspection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_execute_ad_hoc_queryA
Executes an arbitrary SQL query provided by the user via Redshift Data API.
Designed as an escape hatch for advanced users or queries not covered by
specialized tools. Returns a structured dictionary indicating success
(with results) or failure (with error details).
Args:
ctx: The MCP context object.
sql_query: The exact SQL query string to execute.
Returns:
A dictionary conforming to ExecuteAdHocQueryResult structure:
- On success: {"status": "success", "columns": [...], "rows": [...], "row_count": ...}
- On error: {"status": "error", "error_message": "...", "error_type": "..."}
(Note: Actual return might be handled by FastMCP error handling for raised exceptions)
Raises:
DataApiConfigError: If configuration is invalid.
SqlExecutionError: If the SQL execution itself fails.
DataApiTimeoutError: If the Data API call times out.
DataApiError: For other Data API related errors or unexpected issues.
ClientError: For AWS client-side errors.
| Name | Required | Description | Default |
|---|---|---|---|
| sql_query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does an excellent job disclosing behavioral traits. It describes the return structure in detail (success vs error cases), mentions potential exceptions raised (DataApiConfigError, SqlExecutionError, etc.), and notes that 'Actual return might be handled by FastMCP error handling for raised exceptions.' This provides comprehensive behavioral context beyond basic functionality.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and appropriately sized. It begins with the core purpose, then provides usage context, followed by parameter documentation, return value details, and exception information. Every section adds value, though the detailed exception list could be slightly condensed. Overall, it's efficiently organized with clear sections.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (executing arbitrary SQL queries via Redshift Data API) and the absence of both annotations and output schema, the description provides substantial context. It covers purpose, usage guidelines, parameter semantics, return structure, and potential exceptions. The main gap is lack of information about query limitations, performance implications, or security considerations for arbitrary SQL execution.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant meaning beyond the input schema. With 0% schema description coverage (schema only shows sql_query is a required string), the description explains that 'sql_query: The exact SQL query string to execute.' This clarifies the parameter's purpose and format. While it doesn't provide SQL syntax guidance, it adequately compensates for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Executes an arbitrary SQL query provided by the user via Redshift Data API.' It specifies the exact action (execute SQL query), the mechanism (Redshift Data API), and distinguishes it from specialized tools by calling it an 'escape hatch for advanced users or queries not covered by specialized tools.' This differentiates it from sibling tools like handle_get_table_definition or handle_inspect_table.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: 'Designed as an escape hatch for advanced users or queries not covered by specialized tools.' This provides clear guidance that this tool should be used when other specialized tools (the siblings listed) don't cover the needed functionality, establishing clear alternatives and usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_get_table_definitionA
Retrieves the DDL (Data Definition Language) statement for a specific table.
Executes a SQL script designed to generate or retrieve the CREATE TABLE
statement for the given table.
Args:
ctx: The MCP context object.
schema_name: The schema name of the table.
table_name: The name of the table.
Returns:
A dictionary conforming to GetTableDefinitionResult structure:
- On success: {"status": "success", "ddl": "<CREATE TABLE statement>"}
- On table not found or DDL retrieval error:
{"status": "error", "error_message": "...", "error_type": "..."}
Raises:
TableNotFound: If the specified table is not found.
DataApiError: If a critical, unexpected error occurs during execution.
| Name | Required | Description | Default |
|---|---|---|---|
| schema_name | Yes | ||
| table_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does well by detailing success/error return structures, specific exception types (TableNotFound, DataApiError), and the SQL script execution behavior. However, it doesn't mention performance characteristics, rate limits, or authentication requirements that would be helpful for a database tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, execution details, Args, Returns, Raises) and every sentence adds value. It's appropriately sized for a tool with 2 parameters and complex return behavior, with no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 2 parameters, no annotations, and no output schema, the description provides excellent coverage of parameters, return values, and exceptions. The main gap is lack of guidance on when to use versus sibling tools, but otherwise it's quite complete for the tool's complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides explicit parameter documentation in the Args section, clearly explaining what schema_name and table_name represent. With 0% schema description coverage, this comprehensive parameter documentation fully compensates and adds significant value beyond the bare input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Retrieves the DDL statement') and resource ('for a specific table'), distinguishing it from sibling tools like handle_execute_ad_hoc_query or handle_inspect_table. It explicitly mentions the SQL script execution aspect, providing precise functional context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when needing table DDL, but doesn't explicitly state when to use this tool versus alternatives like handle_inspect_table or handle_execute_ad_hoc_query. No guidance is provided on prerequisites, error handling expectations, or specific scenarios where this tool is preferred over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_inspect_tableA
Retrieves detailed information about a specific Redshift table.
Fetches table OID, then concurrently executes various inspection scripts
covering design, storage, health, usage, and encoding.
Args:
ctx: The MCP context object.
schema_name: The schema name of the table.
table_name: The name of the table.
Returns:
A dictionary where keys are script names and values are either the raw
list of dictionary results from the SQL query, the extracted DDL string,
or an Exception object if that specific script failed.
- On success: Dictionary containing raw results or Exception objects for each script.
- On table not found: Raises TableNotFound exception.
- On critical errors (e.g., OID lookup failure): Raises DataApiError or similar.
Raises:
DataApiError: If a critical error occurs during script execution.
TableNotFound: If the specified table cannot be found via its OID.
| Name | Required | Description | Default |
|---|---|---|---|
| schema_name | Yes | ||
| table_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does so effectively. It discloses the concurrent execution of multiple scripts, the mixed return types (raw results, DDL strings, or Exception objects), and specific error conditions (TableNotFound, DataApiError). However, it omits details like rate limits, authentication needs, or performance implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, Args, Returns, Raises) and front-loaded key information. It avoids redundancy, but the Returns section is slightly verbose in detailing success/error cases; some details could be condensed without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description provides substantial context: purpose, parameters, return structure, and error handling. It adequately covers the tool's complexity (2 params, mixed outputs). However, it lacks examples of return values or script names, which would enhance completeness for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly lists and explains the two parameters (schema_name and table_name) in the Args section, clarifying their roles in identifying the Redshift table. This adds meaningful context beyond the bare schema, though it could elaborate on format constraints (e.g., case sensitivity).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Retrieves detailed information') and resource ('about a specific Redshift table'), distinguishing it from siblings like handle_get_table_definition (which likely fetches only DDL) and handle_diagnose_query_performance (which focuses on queries rather than table metadata). The mention of 'various inspection scripts covering design, storage, health, usage, and encoding' provides concrete scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly suggests usage when detailed table metadata is needed, but lacks explicit guidance on when to choose this over alternatives like handle_get_table_definition or handle_monitor_workload. It does not specify prerequisites or exclusions, though the error conditions hint at when-not scenarios (e.g., table not found).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_monitor_workloadA
Analyzes cluster workload patterns over a specified time window.
Executes various SQL scripts concurrently to gather data on resource usage,
WLM performance, top queries, queuing, COPY performance, and disk-based
queries. Returns a dictionary containing the raw results (or Exceptions)
keyed by the script name.
Args:
ctx: The MCP context object.
time_window_days: Lookback period in days for the workload analysis.
Defaults to 2.
top_n_queries: Number of top queries (by total execution time) to
consider for the 'top_queries.sql' script. Defaults to 10.
Returns:
A dictionary where keys are script names (e.g., 'workload/top_queries.sql')
and values are either a list of result rows (as dictionaries) or the
Exception object if that script failed.
Raises:
DataApiError: If a critical error occurs during configuration loading.
(Note: Individual script errors are returned in the result dict).
| Name | Required | Description | Default |
|---|---|---|---|
| time_window_days | No | ||
| top_n_queries | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes that the tool executes SQL scripts concurrently, returns a dictionary with raw results or exceptions, and handles individual script failures gracefully by including exceptions in the result dict. It also mentions that critical configuration errors raise DataApiError. However, it doesn't specify performance characteristics, rate limits, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, execution details, args, returns, raises) and front-loaded with the core functionality. While comprehensive, some sentences could be more concise, such as the detailed explanation of the return dictionary structure which is somewhat verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a workload analysis tool with 2 parameters, no annotations, and no output schema, the description provides substantial context about behavior, parameters, return format, and error handling. It explains the concurrent execution of SQL scripts, the dictionary return structure with success/failure results, and different error scenarios. The main gap is lack of information about what specific workload metrics are analyzed beyond the general categories mentioned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides excellent parameter semantics beyond the basic schema. While schema description coverage is 0%, the description clearly explains that time_window_days is the 'lookback period in days for workload analysis' with a default of 2, and top_n_queries determines 'number of top queries to consider' with a default of 10. This adds meaningful context about what these parameters control in the analysis.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'analyzes cluster workload patterns over a specified time window' with specific verbs (analyzes, executes, gathers) and resources (cluster workload, SQL scripts). It distinguishes from siblings like handle_check_cluster_health or handle_diagnose_query_performance by focusing on comprehensive workload analysis rather than specific health checks or query diagnostics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for analyzing workload patterns over time, but doesn't explicitly state when to use this tool versus alternatives like handle_diagnose_query_performance or handle_execute_ad_hoc_query. There's no guidance on prerequisites, exclusions, or specific scenarios where this tool is preferred over sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.0.0- First observed
handle_check_cluster_health - First observed
handle_diagnose_locks - First observed
handle_diagnose_query_performance - First observed
handle_execute_ad_hoc_query - First observed
handle_get_table_definition - First observed
handle_inspect_table - First observed
handle_monitor_workload
TDQS
Scored across 7 tools
Each tool has a distinct purpose with clear boundaries: cluster health assessment, lock diagnosis, query performance analysis, ad-hoc query execution, table definition retrieval, table inspection, and workload monitoring. There is no functional overlap between tools, and the descriptions clearly differentiate their specific use cases.
All tools follow a 'handle_verb_noun' prefix pattern, which provides some consistency. However, the verb choices are mixed ('check', 'diagnose', 'execute', 'get', 'inspect', 'monitor'), making the naming somewhat inconsistent in terms of action semantics. The structure is predictable but the verb selection lacks uniformity.
With 7 tools, this server is well-scoped for Redshift cluster diagnostics and management. Each tool serves a specific, valuable function in the domain, and there are no redundant or trivial tools. The count is appropriate for covering key operational and troubleshooting tasks without being overwhelming.
The toolset covers essential diagnostic and operational areas for Redshift: health checks, lock analysis, query performance, ad-hoc queries, table definitions, table inspection, and workload monitoring. Minor gaps exist, such as lack of tools for cluster configuration changes, user/role management, or backup operations, but core diagnostic workflows are well-covered.
Maintenance
Related MCP Connectors
Hosted Amazon Seller and Vendor MCP server for Claude, ChatGPT, Cursor, Codex, Gemini, Copilot.
Hosted Amazon Seller Central and Amazon Ads MCP server for Claude, ChatGPT, Cursor, and agents.
MCP server that lets AI assistants use all OneSchema features exposed via the public API.
MCP server for OpenAI API (chat completions, image generation, embeddings) via AceDataCloud
Related MCP Servers
- AlicenseBqualityAmaintenanceModel Context Protocol (MCP) server that integrates Redash with AI assistants like Claude, allowing them to query data, manage visualizations, and interact with dashboards through natural language.4674,035 npm105MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables AI assistants to interact with Amazon Redshift databases, allowing for schema exploration, query execution, and statistics collection.32Apache 2.0
- AlicenseAqualityAmaintenanceA Snowflake MCP server — SQL queries, schema exploration, and data insights for AI assistants62MIT
- AlicenseAqualityCmaintenanceModel Context Protocol (MCP) server for Redash - manage queries, dashboards, and visualizations through AI assistants like Claude.6MIT