Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps with Open Source Solutions

  • Key concepts and strategic advantages of AIOps.
  • The role of Prometheus and Grafana within the observability ecosystem.
  • Positioning ML in AIOps: contrasting predictive and reactive analytics.

Initializing Prometheus and Grafana

  • Installation and configuration of Prometheus for time series data collection.
  • Developing Grafana dashboards utilizing real-time metrics.
  • Investigating exporters, relabeling mechanisms, and service discovery.

Preparing Data for Machine Learning

  • Extraction and transformation of Prometheus metrics.
  • Dataset preparation for anomaly detection and forecasting tasks.
  • Utilizing Grafana’s transformation features or Python-based data pipelines.

Machine Learning for Anomaly Detection

  • Foundational ML models for outlier identification (e.g., Isolation Forest, One-Class SVM).
  • Model training and evaluation using time series data.
  • Visualization of detected anomalies within Grafana dashboards.

Metric Forecasting with Machine Learning

  • Construction of forecasting models (ARIMA, Prophet, and introductory LSTM concepts).
  • Predicting system loads and resource consumption patterns.
  • Leveraging predictions for proactive alerting and scaling decisions.

Integrating ML with Alerting and Automation

  • Formulating alert rules based on ML outputs or defined thresholds.
  • Configuring Alertmanager and notification routing strategies.
  • Automating workflows or script execution in response to detected anomalies.

Scaling and Operationalizing AIOps

  • Integration with external observability platforms (e.g., ELK stack, Moogsoft, Dynatrace).
  • Deploying ML models within observability pipelines.
  • Best practices for implementing AIOps at scale.

Conclusion and Future Directions

Requirements

  • Solid grasp of system monitoring and observability fundamentals.
  • Prior hands-on experience with Grafana or Prometheus.
  • Proficiency in Python and an understanding of core machine learning concepts.

Target Audience

  • Observability engineers.
  • Infrastructure and DevOps teams.
  • Monitoring platform architects and Site Reliability Engineers (SREs).

Number of participants


Price per participant

Upcoming Courses

Related Categories