Buch, Englisch, 360 Seiten, Format (B × H): 156 mm x 234 mm
Buch, Englisch, 360 Seiten, Format (B × H): 156 mm x 234 mm
ISBN: 978-1-041-30679-5
Verlag: Taylor & Francis Ltd
This book provides a detailed guide to programming Graphics Processing Units (GPUs) for high-performance computing, covering both foundational concepts and advanced techniques. It introduces GPU architectures, programming frameworks like CUDA and OpenCL, and optimization strategies for memory and thread management. Readers will learn to implement parallel algorithms, integrate GPU programming with libraries like cuBLAS and TensorRT, and explore advanced topics such as multi-GPU programming and unified memory.
- Explains GPU execution models (SIMD, SIMT), memory hierarchies (global, shared, constant), and thread-level parallelism with practical illustrations.
- Guides through writing kernels, managing grids and threads, synchronization techniques, and error handling in CUDA and OpenCL.
- Covers GPU memory allocation, coalesced memory access, latency hiding, and bandwidth maximization for performance efficiency.
- Provides coding examples for reduction, scan, sorting, and matrix operations, analyzing synchronization, scalability, and performance bottlenecks.
- Demonstrates GPU acceleration in AI and data science workflows using cuBLAS, TensorRT, RAPIDS, PyCUDA, and Numba.
- Explores multi-GPU programming, unified memory, dynamic parallelism, and GPU virtualization, addressing challenges in energy efficiency and portability.
Featuring comprehensive real-world case studies in artificial intelligence, data science, and scientific computing, this resource provides a rigorous and practical foundation for understanding and leveraging GPU acceleration in computationally intensive applications. This book is ideal for students, researchers, and professionals seeking to harness GPU acceleration for computational tasks.
Zielgruppe
Undergraduate Core
Autoren/Hrsg.
Fachgebiete
- Mathematik | Informatik EDV | Informatik Programmierung | Softwareentwicklung Algorithmen & Datenstrukturen
- Mathematik | Informatik EDV | Informatik Programmierung | Softwareentwicklung Spiele-Programmierung, Rendering, Animation
- Mathematik | Informatik EDV | Informatik Programmierung | Softwareentwicklung Programmierung: Methoden und Allgemeines
- Mathematik | Informatik EDV | Informatik Programmierung | Softwareentwicklung Webprogrammierung
- Mathematik | Informatik EDV | Informatik Programmierung | Softwareentwicklung Compiler
Weitere Infos & Material
Chapter 1. Evolution of Parallel Computing. 1.1 The Need for High-Performance Computing. 1.2 From CPUs to GPUs: Architectural Shifts. 1.3 Applications Driving GPU Programming. Chapter 2. Fundamentals of GPU Architecture. 2.1 GPU vs CPU: Key Differences. 2.2 SIMD, SIMT, and Thread-Level Parallelism. 2.3 Memory Hierarchy in GPUs. 2.4 Execution Model and Warp Scheduling. Chapter 3. Getting Started with GPU Programming. 3.1 Introduction to CUDA and OpenCL. 3.2 GPU Programming Workflow. 3.3 Environment and Installing. 3.4 Hello GPU: First GPU Program. Chapter 4. CUDA Programming Essentials. 4.1 Kernel Functions and Thread Hierarchy. 4.2 Grids, Blocks, and Threads. 4.3 Synchronization and Barriers. 4.4 Error Handling in CUDA. Chapter 5. Memory Management in GPUs. 5.1 Types of GPU Memory: Global, Shared, Local, Constant. 5.2 Memory Allocation and Transfer. 5.3 Coalesced Memory Access. 5.4 Optimization Strategies. 5.5 GPU Cache Architecture. Chapter 6. Performance Optimization Techniques in GPU. 6.1 Occupancy and Thread Divergence. 6.2 Shared Memory Utilization. 6.3 Latency Hiding and Instruction Pipelining. 6.4 Profiling GPU Programs. 6.5 Example 1: Using Shared Memory for Faster Access. Chapter 7. Parallel Algorithms on GPUs. 7.1 Introduction Parallel Reduction. 7.2 Scan (Prefix Sum) Algorithms. 7.3 Sorting on GPUs. 7.4 Matrix Operations and Linear Algebra Kernels. 7.5 Example 1: Parallel Reduction (Sum). Chapter 8. OpenCL Programming Model. 8.1 Introduction OpenCL Architecture and Execution Model. 8.2 Kernels and Work-Items. 8.3 Memory Objects and Buffers. 8.4 Comparing CUDA and OpenCL. 8.5 Example 1: Simple OpenCL Kernel. Chapter 9. GPU Libraries and Frameworks. 9.1 cuBLAS and cuFFT. 9.2 Thrust Library for Parallel Algorithms. 9.3 TensorRT and GPU-Accelerated AI Libraries. 9.4 Cross-Vendor Libraries and Portable Frameworks. 9.5 Integration with Python (PyCUDA, Numba). Chapter 10. GPU Programming for Data Science. 10.1 GPU Acceleration in Data Analytics. 10.2 RAPIDS Framework. 10.3 GPU-Accelerated Machine Learning. 10.4 Deep Learning with GPUs. 10.5 Bias Detection and Fairness in AI Lending Decisions derived from NLP Sentiment. Chapter 11. Advancements in GPU Programming. 11.1 Multi-GPU Programming. 11.2 Unified Memory and Heterogeneous Computing. 11.3 Dynamic Parallelism. 11.4 GPU Virtualization and Cloud GPUs. Chapter 12. OpenMP Offloading Architecture and Execution Model. 12.1 How the OpenMP Target Model Maps Work to GPU Devices. 12.2 Overview of Teams, Threads, and SIMD Constructs. 12.3 Interaction with Vendor Backends (NVIDIA, AMD, Intel, ARM). 12.4 Compilation Pipeline: Clang/LLVM, GCC, Intel oneAPI Tooling. 12.5 Example 1: OpenMP Target Offload to GPU. Chapter 13. Case Studies and Applications. 13.1 Scientific Simulations on GPUs. 13.2 Image and Video Processing. 13.3 Cryptography and Blockchain Acceleration. 13.4 Real-Time Systems and Gaming Engines. 13.5 Example 1: GPU Image Smoothing (3×3 Filter). Chapter 14. Debugging and Profiling Tools. 14.1 NVIDIA Nsight and Visual Profiler. 14.2 Performance Counters and Tracing. 14.3 Debugging CUDA Applications. 14.4 Benchmarking Best Practices. 14.5 Example 1: Adding NVTX Annotations for Profiling. Chapter 15. Challenges and Future of GPU Programming. 15.1 Power Consumption and Energy Efficiency. 15.2 Portability Across GPU Architectures. 15.3 GPUs vs TPUs and Other Accelerators. 15.4 Future Trends in GPU Computing. 15.5 Example 1: Measuring GPU Power Using NVML. Chapter 16. Multi-GPU, Multi-Node, and Cloud-Native GPU Execution. 16.1 NCCL, RCCL, and Distributed Communication Topologies. 16.2 Pipeline Parallelism, ZeRO, and Sharded Training Strategies. 16.3 Kubernetes GPU Orchestration, MIG, MPS, and Virtualized GPUs. 16.4 Cloud GPU Computing Workflows on AWS/GCP/Azure. 16.5 Performance Engineering in Large-Scale Distributed Environments. 16.6 NCCL All-Reduce. Chapter 17. GPU Acceleration in Emerging Computational Domains. 17.1 GPU Acceleration in Computer Vision and Imaging Pipelines. 17.2 AI/ML Pipelines Accelerated by GPUs. 17.3 Blockchain and Cryptography on GPUs. 17.4 Quantum Computing Simulation on GPUs.




