This intensive three-day workshop is dedicated to constructing and refining high-performance data-processing workloads using PySpark, Pandas, and Polars within Kubernetes environments.
Learners will gain a hands-on grasp of Spark execution mechanics on Kubernetes, discovering how specific configuration choices impact performance, scalability, resource efficiency, and operational costs. The curriculum delves into critical optimisation domains, such as executor sizing, memory management, dynamic allocation, partitioning logic, shuffle operations, mitigating small-file issues, and enhancing Parquet processing efficiency.
The training also tackles frequent obstacles associated with Pandas, such as memory constraints and out-of-memory crashes, while presenting Polars as a robust, high-performance solution for specific data tasks. Through practical exercises, attendees will learn to diagnose performance and memory bottlenecks, evaluate various configuration strategies, and implement optimisation techniques in realistic ETL and machine learning contexts.
Central to this course is practical decision-making: mastering the identification of bottlenecks, selecting the optimal tool, configuring Spark effectively, and balancing high performance against infrastructure resource usage and cost.
Read more...