Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open-Source Tools
- Overview of AIOps concepts and their strategic benefits
- The role of Prometheus and Grafana within the observability stack
- Positioning ML in AIOps: contrasting predictive and reactive analytics
Configuring Prometheus and Grafana
- Installation and configuration of Prometheus for time series data collection
- Developing Grafana dashboards utilizing real-time metrics
- Investigating exporters, relabeling, and service discovery mechanisms
Data Preprocessing for Machine Learning
- Extracting and transforming Prometheus metrics for analysis
- Structuring datasets for anomaly detection and forecasting models
- Leveraging Grafana transformations or Python-based pipelines
Applying Machine Learning for Anomaly Detection
- Implementing foundational ML models for outlier detection (e.g., Isolation Forest, One-Class SVM)
- Training and assessing models using time series data
- Displaying detected anomalies within Grafana dashboards
Forecasting Metrics with Machine Learning
- Developing basic forecasting models (introduction to ARIMA, Prophet, and LSTM)
- Anticipating system load and resource consumption
- Utilizing predictions for early warning alerts and scaling decisions
Integrating ML with Alerting and Automation
- Establishing alert rules based on ML outputs or predefined thresholds
- Utilizing Alertmanager and configuring notification routing
- Initiating scripts or automation workflows upon anomaly detection
Scaling and Operationalizing AIOps
- Connecting with external observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
- Operationalizing ML models within observability pipelines
- Best practices for implementing AIOps at scale
Summary and Future Steps
Requirements
- A solid grasp of system monitoring and observability principles
- Practical experience with Grafana or Prometheus
- Knowledge of Python and fundamental machine learning concepts
Target Audience
- Observability engineers
- Infrastructure and DevOps teams
- Monitoring platform architects and site reliability engineers (SREs)