Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Performance Concepts and Metrics
- Latency, throughput, power consumption, and resource utilization.
- Distinguishing system-level versus model-level bottlenecks.
- Profiling strategies for inference versus training.
Profiling on Huawei Ascend
- Utilizing CANN Profiler and MindInsight.
- Kernel and operator diagnostics.
- Analyzing offload patterns and memory mapping.
Profiling on Biren GPU
- Leveraging Biren SDK performance monitoring features.
- Optimizing kernel fusion, memory alignment, and execution queues.
- Power and temperature-aware profiling techniques.
Profiling on Cambricon MLU
- Using BANGPy and Neuware performance tools.
- Gaining kernel-level visibility and interpreting logs.
- Integrating the MLU profiler with deployment frameworks.
Graph and Model-Level Optimization
- Strategies for graph pruning and quantization.
- Operator fusion and computational graph restructuring.
- Standardizing input sizes and batch tuning.
Memory and Kernel Optimization
- Optimizing memory layout and reuse patterns.
- Managing buffers efficiently across different chipsets.
- Applying platform-specific kernel tuning techniques.
Cross-Platform Best Practices
- Achieving performance portability through abstraction strategies.
- Developing shared tuning pipelines for multi-chip environments.
- Case study: Tuning an object detection model across Ascend, Biren, and MLU.
Summary and Next Steps
Requirements
- Experience with AI model training or deployment pipelines.
- Understanding of GPU/MLU compute principles and model optimization techniques.
- Basic familiarity with performance profiling tools and metrics.
Audience
- Performance engineers.
- Machine learning infrastructure teams.
- AI system architects.
21 Hours