Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Objectives, and Migration Strategy
- Defining course goals, aligning participant profiles, and establishing success criteria
- Reviewing high-level migration approaches and associated risk factors
- Configuring workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Exploring Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Understanding SMP vs MPP differences and their implications for migration
- Designing the Medallion (Bronze→Silver→Gold) architecture and introducing Unity Catalog
Day 1 Lab — Translating a Stored Procedure
- Practically migrating a sample stored procedure to a notebook
- Mapping temp tables and cursors to DataFrame transformations
- Validating and comparing results with the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- Understanding ACID transactions, commit logs, versioning, and time travel features
- Implementing Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Applying OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning techniques
Day 2 Lab — Incremental Ingestion & Optimization
- Setting up Auto Loader ingestion and MERGE workflows
- Executing OPTIMIZE, Z-ORDER, and VACUUM operations and validating outcomes
- Assessing improvements in read/write performance
Day 3 — SQL in Databricks, Performance & Debugging
- Utilizing analytical SQL features: window functions, higher-order functions, and JSON/array handling
- Interpreting the Spark UI, DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
- Applying query tuning patterns: broadcast joins, hints, caching, and spill reduction
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Leveraging Spark UI traces to identify and resolve skew and shuffle issues
- Benchmarking before and after changes and documenting tuning procedures
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Understanding the Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Converting a procedural ETL script into modular PySpark notebooks
- Incorporating parametrization, unit-style tests, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Designing Databricks Workflows: job structure, task dependencies, triggers, and error handling
- Building incremental Medallion pipelines with quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated with Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Applying Unity Catalog governance, lineage, and access control best practices
- Managing cost, cluster sizing, autoscaling, and job concurrency patterns
- Preparing deployment checklists, rollback strategies, and runbooks
Final Review, Knowledge Transfer, and Next Steps
- Presenting participant migration work and key lessons learned
- Conducting gap analysis, recommending follow-up activities, and handing over training materials
- Providing references, further learning paths, and support options
Requirements
- A solid grasp of data engineering concepts
- Proficiency in SQL and stored procedures (Synapse / SQL Server)
- Knowledge of ETL orchestration concepts (ADF or similar tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers transitioning procedural OLAP logic to Lakehouse patterns
- Platform engineers responsible for overseeing Databricks adoption