Get in Touch
 Duration 14 hours

Course Outline

Readying Machine Learning Models for Production Deployment

  • Encapsulating models using Docker
  • Model export processes from TensorFlow and PyTorch
  • Strategies for versioning and model storage

Serving Models via Kubernetes

  • Introduction to inference server architectures
  • Implementation of TensorFlow Serving and TorchServe
  • Configuration of model service endpoints

Optimizing Inference Performance

  • Effective batching methodologies
  • Management of concurrent requests
  • Tuning for optimal latency and throughput

Dynamic Scaling for ML Workloads

  • Utilization of the Horizontal Pod Autoscaler (HPA)
  • Application of the Vertical Pod Autoscaler (VPA)
  • Event-driven scaling via Kubernetes Event-Driven Autoscaling (KEDA)

GPU Allocation and Resource Governance

  • Setup and configuration of GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Strategies for Model Rollout and Release

  • Blue/green deployment techniques
  • Application of canary rollout patterns
  • A/B testing frameworks for model validation

Production Monitoring and Observability for ML

  • Key metrics for inference workloads
  • Best practices in logging and distributed tracing
  • Designing dashboards and alerting systems

Ensuring Security and System Reliability

  • Security measures for model endpoints
  • Network policies and access control mechanisms
  • Maintaining high availability standards

Conclusion and Future Directions

Requirements

  • A solid grasp of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Basic proficiency in Kubernetes concepts

Target Audience

  • ML Engineers
  • DevOps Engineers
  • Platform Engineering Teams

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories