Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
logging
{}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
extensions
{
  "io.modelcontextprotocol/ui": {}
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
convertB

Convert SQL to PySpark code or process SQL files in batch.

Modes:

sql Convert a single SQL query to PySpark code. Parameters: sql_query (required), table_info, dialect, optimization_level, include_glue_template, style (production | notebook), target (spark | glue)

batch_files Process multiple SQL files into PySpark. Parameters: file_paths (required list), output_dir, job_name

batch_dir Process all SQL files in a directory. Parameters: directory_path (required), output_dir, recursive, job_name

from_pdf Extract SQL from a PDF file and convert to PySpark. Parameters: pdf_path (required)

analyzeA

Deprecated. Prefer convert / glue_job / review. Still registered this minor version.

Analyze SQL or PySpark code for context, data flow, or optimization opportunities.

Modes:

sql_context Analyze SQL context (schemas, tables, dialect, complexity). Parameters: sql_content or selected_text

data_flow Analyze data flow patterns in PySpark code. Parameters: pyspark_code (required), table_info

codebase Analyze a PySpark codebase directory for patterns and issues. Parameters: directory_path (required), include_optimization_suggestions, scan_depth

workspace Full workspace analysis including project structure. Parameters: sql_content or workspace_path, include_project_structure, workspace_name

optimizeA

Deprecated. Prefer review. Still registered this minor version.

Optimize PySpark code and recommend performance improvements.

Modes:

code Return pattern-based suggestions for PySpark code. Does not rewrite the input. Parameters: code (required), optimization_level

joins Recommend join strategies based on estimated table sizes. Parameters: pyspark_code (required), table_info

partitioning Suggest optimal partitioning strategies. Parameters: pyspark_code (required), table_info

comprehensive Generate comprehensive optimization recommendations + performance estimates. Parameters: pyspark_code (required), table_info

reviewA

Review PySpark code for issues, patterns, and refactoring opportunities.

Modes:

code Review PySpark code for issues, best practices, and performance. Parameters: code (required), focus_areas

patterns Analyze code samples to discover common patterns. Parameters: code_samples (required list)

duplicates Detect duplicate patterns across code samples. Parameters: code_samples (required list)

glue_jobA

Generate and manage AWS Glue job configurations and templates.

Modes:

template Generate a complete AWS Glue job template. Parameters: sql_query, job_name, source_database, source_table, target_database, target_table, output_dir, source_format, target_format, include_bookmarking, template_type, script_name

dynamic_frame Convert PySpark DataFrame code to use DynamicFrames. Parameters: pyspark_code (required), source_database, source_table, target_database, target_table

properties Generate Glue job properties for AWS CLI/SDK/Terraform. Parameters: job_name (required), job_type, worker_type, number_of_workers, max_retries, timeout, glue_version, enable_continuous_logging, enable_metrics, enable_spark_ui

sql_conversion Generate a Glue job that includes SQL-to-PySpark conversion. Parameters: sql_query (required), job_name (required), source_database, source_table, target_database, target_table, source_format, target_format, include_bookmarking

glue_schemaA

Deprecated. Prefer glue_job. Still registered this minor version.

Manage Glue Data Catalog schemas — detect, evolve, and define.

Modes:

detect Detect schema from sample data and generate table definition. Parameters: sample_data (required dict/list), table_name (required), infer_partitions

evolve Generate schema evolution strategy for handling schema changes. Parameters: current_columns (required), new_columns (required), merge_behavior, case_sensitive

catalog Generate AWS Glue Data Catalog table definition. Parameters: database_name (required), table_name (required), s3_location (required), data_format, columns, partition_keys, enable_schema_evolution

glue_s3A

Deprecated. Prefer glue_job. Still registered this minor version.

Analyze S3 data layouts using path heuristics (no AWS API call).

Modes:

analyze Path-heuristic layout suggestions. No AWS call; figures are not measured. Parameters: s3_location (required), database_name (required), table_name (required), data_format, query_patterns, data_size_gb

optimize Generate comprehensive S3 optimization strategy. Parameters: database_name (required), table_name (required), s3_location (required), data_format, target_file_size_mb, compression_type, enable_small_file_optimization, query_patterns

consolidate Generate Glue job for small files consolidation. Parameters: source_database (required), source_table (required), target_database (required), target_table (required), target_file_size_mb, consolidation_strategy

glue_dataA

Deprecated. Prefer glue_job. Still registered this minor version.

Generate AWS Glue data processing jobs — incremental, CDC, bookmarks.

Modes:

incremental Generate Glue job with incremental processing and job bookmarking. Parameters: source_database (required), source_table (required), target_database (required), target_table (required), incremental_column (required), incremental_strategy, transformation_sql

cdc Generate Change Data Capture (CDC) Glue job. Parameters: source_database (required), source_table (required), target_database (required), target_table (required), cdc_column, cdc_strategy, primary_keys

bookmarks Generate job bookmark configuration for Glue jobs. Parameters: job_name (required), bookmark_strategy, transformation_context_keys

refactorA

Deprecated. Prefer review. Still registered this minor version.

Refactor PySpark code and generate pipeline structures.

Modes:

patterns Refactor code by replacing duplicate patterns with utility function calls. Parameters: original_code (required), code_samples (required)

utilities Extract common utility functions from code patterns. Parameters: code_samples (required), patterns

pipeline Generate optimized PySpark data pipeline code or project structure. Parameters (pipeline): data_sources (required list), processing_requirements (required), target_format, include_monitoring Parameters (project): sql_content, workspace_name, workspace_path, output_dir, include_glue_template, dialect, include_batch_processing, include_visualization

searchA

Deprecated. Prefer convert. Still registered this minor version.

Search stored conversions, code patterns, and context data.

Modes:

conversions Search previously converted SQL queries and history. Parameters: query, limit If query is empty, returns recent conversion history.

patterns Search stored code patterns by description or template. Parameters: query, limit, min_usage_count If query is empty, returns all stored patterns with min_usage_count.

context Retrieve stored conversion context. Parameters: conversion_id or key

contextA

Deprecated. Prefer convert. Still registered this minor version.

Store, retrieve, and work with SQL/PySpark conversion context.

Modes:

store Store additional context data for a conversion. Parameters: conversion_id (required), context_data (required)

get Retrieve stored context for a conversion. Parameters: conversion_id (required)

assist Real-time SQL assistance — analyze and convert as you edit. Parameters: sql_query or selected_text

batch_statusA

Deprecated. Prefer convert(mode="batch_dir"). Still registered this minor version.

Monitor and manage batch processing jobs.

Modes:

status Get the status of a specific batch job. Parameters: job_id (required)

cancel Cancel a running batch job. Parameters: job_id (required)

active List all currently active batch jobs. Parameters: none

recent List recent batch jobs. Parameters: limit, status

s3_sourceA

Deprecated. Prefer glue_job. Still registered this minor version.

Analyze S3 data sources and Delta tables.

Modes:

analyze Analyze S3 data source structure, format, and optimization opportunities. Parameters: s3_path (required), include_schema_inference

delta Analyze Delta table structure, properties, and optimization. Parameters: table_path (required), analyze_history

analyticsA

Deprecated. Prefer review. Still registered this minor version.

Analytics on optimization effectiveness and usage patterns.

Modes:

optimization Get analytics on optimization effectiveness. Parameters: optimization_type, limit

usage Get usage statistics including conversion history and pattern stats. Parameters: limit

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.4/5.0

Scored across 14 tools

Disambiguation2/5

Multiple tools have unclear boundaries: analyze, optimize, review, and refactor all target PySpark code analysis/improvement, while glue_s3, s3_source, glue_job, and glue_data overlap on Glue/S3 concerns. The deprecation notes help steer agents, but the large number of legacy tools still creates significant selection ambiguity.

Naming Consistency2/5

Tool naming mixes single-word verbs (convert, analyze, review, search), noun-style names (context, analytics), and compound names in inconsistent orders (glue_s3 vs s3_source, glue_job vs batch_status). Though all are lowercase snake_case, there is no predictable verb_noun or noun_noun convention across the set.

Tool Count3/5

Fourteen tools is a reasonable raw count, but 11 of them are explicitly deprecated, leaving only three actively preferred tools. The surface is bloated with redundant legacy tools that add noise without expanding genuine capability.

Completeness4/5

The core workflows—SQL-to-PySpark conversion, PySpark review, and Glue job generation—are well covered, with supporting capabilities for schemas, S3 layout analysis, batch processing, and context storage. Minor gaps exist, such as no non-deprecated tool for S3 source analysis or batch job status, but these are workable via the preferred tools.

Maintenance

ActivitySlowing
ResponsivenessNo issues