PySpark MCP Server
Related Servers
Alternatives to PySpark MCP Server
No user-submitted related servers found.
Related Servers
- FlicenseNot gradedqualityDmaintenanceProvides specialized tools for data engineering tasks like SQL formatting, dbt model generation, and Snowflake table creation. It enables users to analyze CSV data, validate pipeline configurations, and summarize ETL lineage through natural language.-
- AlicenseAqualityDmaintenanceProvides SQL analysis, linting, and dialect conversion using SQLGlot, enabling validation, transpilation, and extraction of table/column references.432MIT
- FlicenseNot gradedqualityDmaintenanceEnables natural language-powered ETL workflows using Airflow, AWS Glue, Athena, and S3, allowing LLM agents to control and monitor data infrastructure.-
- AlicenseAqualityDmaintenanceEnables AI assistants to query Apache Iceberg tables on S3 via AWS Glue Data Catalog using DuckDB as the embedded query engine, supporting columnar Arrow reads with no data movement.4MIT
- AlicenseNot gradedqualityCmaintenanceAccess your Databricks workspace through Claude and other LLMs. Query Unity Catalog tables, inspect jobs, and retrieve detailed metadata.7MIT
- FlicenseNot gradedqualityCmaintenanceEnables natural-language queries on Trino big data platforms, generating validated, schema-aware SQL via RAG and local LLM inference, and exposes metadata, query, and profiling tools through MCP.-
TDQS
Scored across 14 tools
Multiple tools have unclear boundaries: analyze, optimize, review, and refactor all target PySpark code analysis/improvement, while glue_s3, s3_source, glue_job, and glue_data overlap on Glue/S3 concerns. The deprecation notes help steer agents, but the large number of legacy tools still creates significant selection ambiguity.
Tool naming mixes single-word verbs (convert, analyze, review, search), noun-style names (context, analytics), and compound names in inconsistent orders (glue_s3 vs s3_source, glue_job vs batch_status). Though all are lowercase snake_case, there is no predictable verb_noun or noun_noun convention across the set.
Fourteen tools is a reasonable raw count, but 11 of them are explicitly deprecated, leaving only three actively preferred tools. The surface is bloated with redundant legacy tools that add noise without expanding genuine capability.
The core workflows—SQL-to-PySpark conversion, PySpark review, and Glue job generation—are well covered, with supporting capabilities for schemas, S3 layout analysis, batch processing, and context storage. Minor gaps exist, such as no non-deprecated tool for S3 source analysis or batch job status, but these are workable via the preferred tools.