Get in Touch
 Duration 35 hours

Course Outline

Introduction, Learning Objectives, and Migration Strategy

  • Course goals, alignment with participant profiles, and definitions of success criteria
  • Overview of high-level migration approaches and associated risk factors
  • Configuration of workspaces, repositories, and lab datasets

Day 1 — Migration Fundamentals and Architecture

  • Lakehouse concepts, Delta Lake overview, and Databricks system architecture
  • Comparative analysis of SMP vs MPP architectures and their impact on migration
  • Design principles of the Medallion (Bronze→Silver→Gold) architecture and an introduction to Unity Catalog

Day 1 Lab — Translating a Stored Procedure

  • Practical exercise: migrating a sample stored procedure into a notebook
  • Converting temporary tables and cursors into DataFrame transformations
  • Validating outputs against the original procedure results

Day 2 — Advanced Delta Lake Features & Incremental Loading

  • Understanding ACID transactions, commit logs, versioning, and time travel capabilities
  • Implementing Auto Loader, MERGE INTO patterns, upserts, and handling schema evolution
  • Storage optimization techniques including OPTIMIZE, VACUUM, Z-ORDER, and partitioning

Day 2 Lab — Incremental Ingestion & Optimization

  • Building Auto Loader ingestion pipelines and MERGE workflows
  • Applying OPTIMIZE, Z-ORDER, and VACUUM commands; verifying outcomes
  • Evaluating improvements in read/write performance metrics

Day 3 — SQL in Databricks, Performance Analysis & Debugging

  • Advanced analytical SQL features: window functions, higher-order functions, and JSON/array manipulation
  • Interpreting Spark UI, DAGs, shuffles, stages, and tasks to diagnose bottlenecks
  • Strategies for query tuning: broadcast joins, hints, caching, and reducing spill

Day 3 Lab — SQL Refactoring & Performance Tuning

  • Refactoring a resource-intensive SQL process into optimized Spark SQL
  • Utilizing Spark UI traces to detect and resolve data skew and shuffle issues
  • Conducting before-and-after benchmarks and documenting tuning procedures

Day 4 — Tactical PySpark: Replacing Procedural Logic

  • Spark execution model deep-dive: drivers, executors, lazy evaluation, and partitioning strategies
  • Converting loops and cursors into efficient, vectorized DataFrame operations
  • Code modularization, UDFs/pandas UDFs, widgets, and creation of reusable libraries

Day 4 Lab — Refactoring Procedural Scripts

  • Converting a procedural ETL script into modular PySpark notebooks
  • Incorporating parametrization, unit-style tests, and reusable functions
  • Conducting code reviews and applying best-practice checklists

Day 5 — Orchestration, End-to-End Pipeline & Best Practices

  • Databricks Workflows: job design, task dependencies, triggers, and error management
  • Architecting incremental Medallion pipelines with quality rules and schema validation
  • Integration with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark

Day 5 Lab — Building a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
  • Implementing logging, auditing, retry mechanisms, and automated validations
  • Executing the full pipeline, verifying outputs, and drafting deployment documentation

Operationalization, Governance, and Production Readiness

  • Best practices for Unity Catalog governance, data lineage, and access control
  • Managing costs, cluster sizing, autoscaling, and job concurrency patterns
  • Creating deployment checklists, rollback strategies, and operational runbooks

Final Review, Knowledge Transfer, and Next Steps

  • Participant presentations showcasing migration work and key learnings
  • Gap analysis, recommended follow-up actions, and handover of training materials
  • Resource references, further learning pathways, and support options

Requirements

  • A solid grasp of core data engineering principles
  • Practical experience with SQL and stored procedures (such as Synapse / SQL Server)
  • Familiarity with ETL orchestration frameworks (like ADF or similar tools)

Target Audience

  • Technology managers possessing a background in data engineering
  • Data engineers looking to transition procedural OLAP logic to Lakehouse paradigms
  • Platform engineers tasked with overseeing the adoption of Databricks

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories