Get in Touch

Course Outline

Introduction to Agentic AI in Operations

  • The shift from static runbooks to reasoning agents: the evolution of IT automation
  • Agent structure: reasoning loops, tool utilisation, memory management, and planning
  • Determining when to automate versus when to retain human oversight

Agent Frameworks and System Architectures

  • Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
  • Multi-agent architectures: supervisor, hierarchical, and swarm models
  • Framework analysis: LangGraph, CrewAI, AutoGen, and custom agent implementations
  • Constructing a basic operational agent: querying monitoring, diagnosing issues, and proposing solutions

Tool Integration for IT Operations

  • Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
  • Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
  • Leveraging infrastructure tools: kubectl, Terraform, and Ansible through agent actions
  • Designing secure tool interfaces with parameter validation and idempotency checks

Automating Incident Response

  • Automated incident triage: severity classification and routing logic
  • Generating root cause hypotheses and collecting supporting evidence
  • Automated remediation tasks: restarting services, scaling resources, rolling back changes, and failovers
  • Creating incident runbook agents with configurable levels of autonomy

Safety, Guardrails, and Human-in-the-Loop Mechanisms

  • Classifying actions: read-only, low-risk, high-risk, and destructive operations
  • Establishing approval gates and escalation policies for critical tasks
  • Guardrail strategies: action allowlists, blast radius limitations, and rollback assurances
  • Maintaining audit trails and decision provenance for regulatory compliance

Multi-Agent Orchestration for Complex Incidents

  • Coordinating specialist agents: triage, diagnosis, and remediation units
  • Managing inter-agent communication and shared contextual data
  • Resolving conflicts when agents suggest opposing actions
  • Simulating major end-to-end incidents with multi-agent response protocols

Observability and Performance Evaluation

  • Tracing agent reasoning chains for debugging and audit purposes
  • Assessing agent decision quality: precision, recall, and time-to-resolution metrics
  • Implementing feedback loops: learning from operator overrides and operational outcomes
  • Tracking costs and managing token economics for operational agents

Production Deployment and Operational Management

  • Deploying agents as services: APIs, webhooks, and scheduled tasks
  • Gradual autonomy rollout: transitioning from shadow mode to full auto-remediation
  • Handling agent failures: operational protocols when the agent encounters errors
  • Constructing the business case and measuring ROI for autonomous operations

Requirements

  • Practical experience with IT operations, DevOps, or SRE workflows.
  • Proficiency in Python scripting and RESTful APIs.
  • Fundamental knowledge of LLM capabilities and prompt engineering techniques.

Target Audience

  • SRE and DevOps professionals exploring AI-driven automation strategies.
  • Platform engineers focused on developing self-healing infrastructure.
  • IT operations leaders evaluating agentic AI solutions for incident management.
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories