Best Apache Spark MCP Servers
Apache Spark is an open-source unified analytics engine for large-scale data processing. It provides high-level APIs in Java, Scala, Python and R, and an optimized engine that supports general execution graphs.
Why this server?
Provides tools for searching and retrieving Apache Spark documentation, enabling full-text keyword searches with section filtering and access to the full content of documentation pages.
AlicenseAqualityAmaintenanceProvides full-text search and retrieval tools for Apache Spark documentation using SQLite FTS5 with BM25 ranking. It enables AI assistants to efficiently search, filter by section, and read specific Spark documentation pages.2MITWhy this server?
Provides read-only access to Apache Spark data through SQL models, allowing for querying live data via natural language questions without requiring SQL knowledge. Tools include listing available tables, retrieving column information, and executing SQL SELECT queries against Spark.
AlicenseNot gradedqualityDmaintenanceThis read-only MCP Server allows you to connect to Apache Spark data from Claude Desktop through CData JDBC Drivers. For full CRUD support, check out the first managed MCP platform: CData Connect AI (https://www.cdata.com/ai/).MITWhy this server?
Offers information about Duyet's expertise with Apache Spark through CV resources and tools, enabling discussions about data engineering projects.
FlicenseBqualityAmaintenanceAn experimental Model Context Protocol server that enables AI assistants to access information about Duyet, including his CV, blog posts, and GitHub activity through natural language queries.82Why this server?
Provides tools for diagnosing Spark job failures and optimizing Spark performance, using error logs and source code to identify root causes and suggest fixes.
AlicenseNot gradedqualityCmaintenanceMCP server that diagnoses Apache Spark job failures and optimizes performance using stack-trace analysis and LLM providers, supporting EMR and local sources.1MITWhy this server?
Provides searchable documentation for Apache Spark as part of the data engineering knowledge base.
AlicenseNot gradedqualityDmaintenanceProvides AI assistants with searchable access to documentation from 170+ curated repositories and 1000+ popular GitHub projects across 20+ categories including trading, AI/ML, DevOps, and web development.3MITWhy this server?
Query Spark SQL clusters via Thrift/HiveServer2 protocol, enabling read-only SQL queries, schema discovery, and multiple authentication methods.
AlicenseNot gradedqualityCmaintenanceAn MCP server that enables AI assistants to query Spark SQL clusters via the Thrift/HiveServer2 protocol.MITWhy this server?
Utilizes Apache Spark for writing Parquet/ORC file formats to MinIO storage as part of data processing pipelines.
AlicenseNot gradedqualityBmaintenanceMCP server with 32 tools for ETL ingestion, AI-generated data quality rules, AI transformations, vector search, and natural-language SQL. Works across Postgres, MongoDB, Kafka, S3/MinIO, HashiCorp Vault, and five vector stores (Qdrant, Weaviate, Milvus, Chroma, pgvector).12AGPL 3.0Why this server?
Supports Apache Spark SQL dialect for SQL generation, schema introspection, validation, and transpilation.
Why this server?
Provides query optimization and data discovery capabilities for Apache Spark by exposing logical and physical query plans, catalog and table information to AI systems.
AlicenseNot gradedqualityCmaintenanceA server implementation of MCP for Apache Spark that provides query plans and catalog information to AI systems for query optimization and data discovery.18Apache 2.0