Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Fundamentals of Cloud Operations on AWS
- Clarifying operational roles and responsibilities within the cloud.
- Structuring AWS accounts, organizations, and multi-account strategies.
- Utilizing core operational services such as CloudWatch, CloudTrail, and AWS Config.
Infrastructure as Code and Provisioning
- Exploring IaC principles and immutable infrastructure.
- Provisioning resources using Terraform and AWS CloudFormation.
- Handling state, modules, and environment promotion workflows.
CI/CD and Deployment Strategies
- Constructing CI/CD pipelines for cloud-native applications.
- Implementing blue/green, canary, and rolling deployment techniques.
- Automating rollbacks, health checks, and release validation processes.
Monitoring, Observability, and Alerting
- Managing metrics, logs, and traces: collection, storage, and analysis.
- Leveraging CloudWatch, X-Ray, and third-party observability tools.
- Establishing SLOs/SLIs, alerting policies, and on-call procedures.
Security Operations and Identity Management
- Applying IAM best practices, least privilege models, and cross-account access control.
- Managing secrets, KMS, and secure parameter stores.
- Enhancing operational security through patching, vulnerability scanning, and audit trails.
Resilience, Backup, and Disaster Recovery
- Engineering for fault tolerance and high availability.
- Defining backup strategies, snapshot automation, and restoration procedures.
- Planning disaster recovery and developing response runbooks.
Cost Optimization and Governance
- Gaining cost visibility through billing, tagging, and allocation strategies.
- Optimizing resources via rightsizing, reserved instances/savings plans, and budget controls.
- Enforcing governance through policies, guardrails, and compliance automation.
Containers, Serverless, and Runtime Operations
- Addressing operational aspects of ECS, EKS, and Lambda.
- Managing service discovery, autoscaling, and resource limitations.
- Logging, tracing, and debugging containerized workloads.
Incident Response, Playbooks, and Chaos Engineering
- Executing runbook-driven incident response and postmortem analysis.
- Automating remediation and implementing self-healing patterns.
- Introduction to chaos experiments for resilience validation.
Practical Workshop: Operating a Sample Workload
- Deploying a sample application using IaC and a CI/CD pipeline.
- Setting up monitoring, alerts, and automated remediation scripts.
- Simulating incidents to practice runbook-based responses.
Summary and Future Directions
Requirements
- Fundamental knowledge of cloud concepts and networking.
- Proficiency with the Linux command line and scripting.
- Familiarity with source control (Git) and foundational CI/CD concepts.
Target Audience
- Cloud operations engineers.
- Site Reliability Engineers (SREs) and platform engineers.
- DevOps engineers and technical team leads.
21 Hours
Testimonials (1)
I've find out new interesting things about Lambda and Serverless