Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Production

  • Encapsulating models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Considerations for version control and storage

Serving Models via Kubernetes

  • Introduction to inference server architectures
  • Deploying TensorFlow Serving and TorchServe
  • Establishing model service endpoints

Techniques for Optimizing Inference

  • Implementing efficient batching strategies
  • Managing concurrent request processing
  • Tuning for optimal latency and throughput

Autoscaling Strategies for ML Workloads

  • Horizontal Pod Autoscaler (HPA) implementation
  • Vertical Pod Autoscaler (VPA) configuration
  • Event-driven autoscaling with Kubernetes Event-Driven Autoscaling (KEDA)

Managing GPU Resources and Allocation

  • Setting up GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML tasks

Strategies for Model Rollout and Release

  • Implementing blue/green deployment patterns
  • Utilizing canary rollout methods
  • Conducting A/B testing for model assessment

Monitoring and Observability in Production

  • Tracking key metrics for inference workloads
  • Best practices for logging and distributed tracing
  • Configuring dashboards and alerting systems

Focus on Security and Reliability

  • Protecting model endpoints
  • Applying network policies and access controls
  • Guaranteeing high availability

Conclusion and Recommended Next Steps

Requirements

  • Knowledge of application workflows in containerized environments
  • Practical experience with Python-based machine learning models
  • Foundational understanding of Kubernetes

Target Audience

  • Machine Learning Engineers
  • DevOps Engineers
  • Platform Engineering Teams

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories