Get in Touch
 Duration 35 hours

Course Outline

Introduction, Objectives, and Migration Strategy

  • Defining course goals, aligning participant profiles, and establishing success criteria
  • Reviewing high-level migration approaches and associated risk factors
  • Configuring workspaces, repositories, and lab datasets

Day 1 — Migration Fundamentals and Architecture

  • Exploring Lakehouse concepts, Delta Lake overview, and Databricks architecture
  • Understanding SMP vs MPP differences and their implications for migration
  • Designing the Medallion (Bronze→Silver→Gold) architecture and introducing Unity Catalog

Day 1 Lab — Translating a Stored Procedure

  • Practically migrating a sample stored procedure to a notebook
  • Mapping temp tables and cursors to DataFrame transformations
  • Validating and comparing results with the original output

Day 2 — Advanced Delta Lake & Incremental Loading

  • Understanding ACID transactions, commit logs, versioning, and time travel features
  • Implementing Auto Loader, MERGE INTO patterns, upserts, and schema evolution
  • Applying OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning techniques

Day 2 Lab — Incremental Ingestion & Optimization

  • Setting up Auto Loader ingestion and MERGE workflows
  • Executing OPTIMIZE, Z-ORDER, and VACUUM operations and validating outcomes
  • Assessing improvements in read/write performance

Day 3 — SQL in Databricks, Performance & Debugging

  • Utilizing analytical SQL features: window functions, higher-order functions, and JSON/array handling
  • Interpreting the Spark UI, DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
  • Applying query tuning patterns: broadcast joins, hints, caching, and spill reduction

Day 3 Lab — SQL Refactoring & Performance Tuning

  • Refactoring complex SQL processes into optimized Spark SQL
  • Leveraging Spark UI traces to identify and resolve skew and shuffle issues
  • Benchmarking before and after changes and documenting tuning procedures

Day 4 — Tactical PySpark: Replacing Procedural Logic

  • Understanding the Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
  • Converting loops and cursors into vectorized DataFrame operations
  • Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries

Day 4 Lab — Refactoring Procedural Scripts

  • Converting a procedural ETL script into modular PySpark notebooks
  • Incorporating parametrization, unit-style tests, and reusable functions
  • Conducting code reviews and applying best-practice checklists

Day 5 — Orchestration, End-to-End Pipeline & Best Practices

  • Designing Databricks Workflows: job structure, task dependencies, triggers, and error handling
  • Building incremental Medallion pipelines with quality rules and schema validation
  • Integrating with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic

Day 5 Lab — Build a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated with Workflows
  • Implementing logging, auditing, retries, and automated validations
  • Executing the full pipeline, validating outputs, and preparing deployment documentation

Operationalization, Governance, and Production Readiness

  • Applying Unity Catalog governance, lineage, and access control best practices
  • Managing cost, cluster sizing, autoscaling, and job concurrency patterns
  • Preparing deployment checklists, rollback strategies, and runbooks

Final Review, Knowledge Transfer, and Next Steps

  • Presenting participant migration work and key lessons learned
  • Conducting gap analysis, recommending follow-up activities, and handing over training materials
  • Providing references, further learning paths, and support options

Requirements

  • A solid grasp of data engineering concepts
  • Proficiency in SQL and stored procedures (Synapse / SQL Server)
  • Knowledge of ETL orchestration concepts (ADF or similar tools)

Target Audience

  • Technology managers with a background in data engineering
  • Data engineers transitioning procedural OLAP logic to Lakehouse patterns
  • Platform engineers responsible for overseeing Databricks adoption

Number of participants


Price per participant

Upcoming Courses

Related Categories