Schulungsübersicht
Introduction to GPU-Accelerated Computing
- Heterogeneous computing and the CPU-GPU architecture
- CUDA execution and memory spaces
- Compiling CUDA C++ with nvcc and CMake
- Verifying the GPU development environment
Parallel Algorithms with Thrust and CUB
- GPU-accelerated sort, reduce, and transform
- Refactoring STL algorithms for GPU execution
- Thrust device vectors and execution policies
- CUB device-wide primitives for custom pipelines
GPU Memory Architecture and Management
- Global, constant, and texture memory types
- Explicit device memory allocation and transfers
- Unified Memory for simplified data access
- Memory coalescing and access pattern optimization
Asynchronous Execution with CUDA Streams
- Creating and managing CUDA streams
- Overlapping kernel execution with data transfers
- CUDA events for dependency management
- Stream priorities and concurrency tuning
Writing Custom CUDA Kernels
- The SIMT programming model and warp execution
- Kernel launch configuration and grid-stride loops
- Thread indexing and multidimensional grids
- Error handling and CUDA runtime API checks
Thread Hierarchy and Execution Model
- Grids, blocks, and threads in device code
- Warp-level primitives and ballot operations
- Block-level synchronization and barriers
- Occupancy and resource utilization analysis
Cooperative Groups for Flexible Parallelism
- The cooperative_groups API and group types
- Thread-block tiles and tiled_partition
- Grid-level cooperative launches
- Multi-grid synchronization patterns
Shared Memory Optimization Techniques
- Shared memory banks and bank conflict avoidance
- Tiling strategies for matrix operations
- Shared memory as user-managed cache
- cuda::shared_memory_mdspan for multidimensional views
Kernel Fusion and Advanced Parallel Patterns
- Fusing multiple kernels to reduce launch overhead
- Scan, reduce-by-key, and segmented algorithms
- Atomic operations and lock-free data structures
- Warp-aggregated atomics for throughput
Profiling and Optimization with Nsight Systems
- Timeline analysis of CPU and GPU activity
- Identifying memory transfer bottlenecks
- Kernel performance and occupancy profiling
- Iterative optimization with Nsight Compute
Modern C++ Features in CUDA Device Code
- Lambdas, constexpr, and auto in kernels
- C++17 parallel algorithms and execution policies
- C++20 concepts and ranges on device
- C++23 support in nvcc and CCCL 3.x
CUDA Graphs and Advanced Asynchrony
- Defining and launching CUDA graphs
- Graph capture from stream execution
- Graph update and conditional execution nodes
- Reducing launch latency for iterative workloads
Integration Patterns for Existing Applications
- Wrapping GPU code behind C++ interfaces
- Managing multi-GPU and NUMA systems
- Build system integration with CMake and CUDA
- Debugging device code with cuda-gdb
Summary and Best Practices
- Choosing between Thrust, CUB, and custom kernels
- Performance portability across GPU architectures
- Code organization and RAII for CUDA resources
- Next steps and advanced CUDA learning paths
Voraussetzungen
- Basis-C++-Kompetenz einschließlich Lambda-Ausdrücke, Templates und der STL
- Vertrautheit mit Standardalgorithmen, Containern und Iteratoren
- Sicherheit im Umgang mit Schleifen, Bedingungen und Funktionen
- Keine Vorkenntnisse in CUDA- oder GPU-Programmierung erforderlich
Zielgruppe
- C++-Entwickler, die rechenintensive Anwendungen mit GPUs beschleunigen möchten
- Softwareingenieure, vom reinen CPU-Betrieb zur heterogenen Parallelprogrammierung wechseln
- Performance-Ingenieure und quantitative Entwickler, die mit großen Datensätzen arbeiten
Erfahrungsberichte (3)
Anfangs wirkte das Tempo des Trainers für mich etwas zu schnell, aber nachdem ich während der Schulung entsprechendes Feedback gegeben hatte, erkannte er dies an und reduzierte das Tempo, ohne dabei an der Qualität der Vorträge zu verlieren. Er baute eine gute Beziehung zum Publikum auf, war sehr freundlich und offen für Diskussionen.
Alexandru Ostafi - Siemens
Kurs - Advanced C++ : Practical workshop
Maschinelle Übersetzung
Detaillierte Erklärungen und subtile Wiederholungen der Punkte, die das Wissen wirklich nachhaltig verankert haben. Rods Bereitschaft, auch selten gestellte Fragen zu überprüfen, um sicherzustellen, dass seine Antworten 100% korrekt waren. Ebenso sein Interesse daran, die Vor- und Nachteile alternativer Programmierstile zu diskutieren, sodass wir nicht nur lernten, wie man C++ in der beabsichtigten Weise verwendet, sondern auch, warum es so gemacht werden sollte.
Nick Dillon - cellxica Ltd
Kurs - Using C++ in Embedded Systems - Applying C++11/C++14
Maschinelle Übersetzung
Erfahrungstechnik, es ist das Wissen und die wertvollen Kenntnisse des Lehrers.
Carey Fan - Logitech
Kurs - C/C++ Secure Coding
Maschinelle Übersetzung