Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- The shift from static runbooks to reasoning agents: the evolution of IT automation
- Agent structure: reasoning loops, tool utilisation, memory management, and planning
- Determining when to automate versus when to retain human oversight
Agent Frameworks and System Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Framework analysis: LangGraph, CrewAI, AutoGen, and custom agent implementations
- Constructing a basic operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: kubectl, Terraform, and Ansible through agent actions
- Designing secure tool interfaces with parameter validation and idempotency checks
Automating Incident Response
- Automated incident triage: severity classification and routing logic
- Generating root cause hypotheses and collecting supporting evidence
- Automated remediation tasks: restarting services, scaling resources, rolling back changes, and failovers
- Creating incident runbook agents with configurable levels of autonomy
Safety, Guardrails, and Human-in-the-Loop Mechanisms
- Classifying actions: read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical tasks
- Guardrail strategies: action allowlists, blast radius limitations, and rollback assurances
- Maintaining audit trails and decision provenance for regulatory compliance
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation units
- Managing inter-agent communication and shared contextual data
- Resolving conflicts when agents suggest opposing actions
- Simulating major end-to-end incidents with multi-agent response protocols
Observability and Performance Evaluation
- Tracing agent reasoning chains for debugging and audit purposes
- Assessing agent decision quality: precision, recall, and time-to-resolution metrics
- Implementing feedback loops: learning from operator overrides and operational outcomes
- Tracking costs and managing token economics for operational agents
Production Deployment and Operational Management
- Deploying agents as services: APIs, webhooks, and scheduled tasks
- Gradual autonomy rollout: transitioning from shadow mode to full auto-remediation
- Handling agent failures: operational protocols when the agent encounters errors
- Constructing the business case and measuring ROI for autonomous operations
Requirements
- Practical experience with IT operations, DevOps, or SRE workflows.
- Proficiency in Python scripting and RESTful APIs.
- Fundamental knowledge of LLM capabilities and prompt engineering techniques.
Target Audience
- SRE and DevOps professionals exploring AI-driven automation strategies.
- Platform engineers focused on developing self-healing infrastructure.
- IT operations leaders evaluating agentic AI solutions for incident management.
14 Hours