Buch, Englisch, 598 Seiten, Format (B × H): 155 mm x 235 mm
17th International Symposium, APPT 2026, Brussels, Belgium, July 27, 2026, Proceedings
Buch, Englisch, 598 Seiten, Format (B × H): 155 mm x 235 mm
Reihe: Lecture Notes in Computer Science
ISBN: 978-981-9248-04-9
Verlag: Springer
This book constitutes the refereed proceedings of the 17th International Symposium on Advanced Parallel Processing Technologies, APPT 2026, held in Brussels, Belgium, on July 27, 2026.
The 36 full papers, 11 Poster Papers and 4 work in- progress papers included in this book were carefully reviewed and selected from 135 submissions. They were organized in topical sections as follows: Featured Articles; LLM Inference; Serving Systems; LLM Workflows; AI Design Automation; Scalable Computing; AI Accelerators; Kernel Acceleration; Streaming Accelerators; Secure Systems; Data Efficiency; Poster and Work-in-Progress Papers.
Zielgruppe
Research
Autoren/Hrsg.
Fachgebiete
Weitere Infos & Material
.- Featured Articles.
.- ZIPPer: A Fine-Grained Co-Design of ZeRO-DP and Pipeline Parallelism for LLM Training on Heterogeneous GPUs.
.- obliv-clang: Real-World Oblivious Programming in C++.
.- Uni-Winograd: A Massively Resource-Efficient and Unified Radix-4 NTT Architec-ture for PQC Algorithms.
.- Barbell: An Extensible Generator of On-Demand Loads for Interactive Cloud Ser-vices.
.- LLM Inference.
.- JSQKV: Joint Sparsification and Quantization for KV-Cache Compression and De-code Acceleration.
.- LocalKV: Leveraging Sparse Attention Locality for Efficient Long-Context LLM Inference.
.- MOCAP: Wafer-Scale-Chip-Oriented Memory-Orchestrated Chunked Pipelining Framework for Prefill-Only LLM Inference.
.- Window-Diffusion: Accelerating Diffusion Language Model Inference with Win-dowed Token Pruning and Caching.
.- Serving Systems.
.- TokenSimV2: Accurate and Fast Modeling for LLM Inference on GPUs.
.- TierServe: Revenue-Maximizing LLM Inference Scheduling across Heterogeneous Subscription Tiers.
.- Identifying Iso-Cost Sweet Spot Configurations for Cloud OLAP Queries.
.- LLM Workflows.
.- Characterizing and Mitigating Context Redundancy in LLM Agent Workflows.
.- Do Not Let Sandboxes Sit Idle: Cross-Agent Sandbox Re-allocation for LLM Agents.
.- CGR: A Budget-Aware Closed-Loop Framework for Efficient Knowledge Graph Question Answerin.
.- AI Design Automation.
.- SysVCoder: An LLM-Driven Framework for Systematic Generation of System-Level Design.
.- TensorAgent: Multi-Agent Framework for Automated Tensor Core Code Generation.
.- CUTE-XS: Bottleneck Analysis and Collaborative Optimization for Integrating CUTE into XiangShan.
.- Scalable Computing.
.- C3FT: Computation-Centric Checkpointing for Distributed Large Model Training.
.- FlexUSFL: A Scalable and Memory-Efficient U-Shaped Split Federated Learning Framework for Large Language Model Fine-Tuning.
.- SAW3D: A Performance-Portable Solver for Compressible Turbulence at the O(10¹°)-Cell Scale on Heterogeneous Supercomputers.
.- AI Accelerators.
.- BitFly: A Low-Bit Mixed-Precision Acceleration Framework for Edge RISC-V Vec-tor Processors.
.- Fold-FP4: Product-Space Folding for Energy-Efficient FP4 Matrix Multiplication.
.- Helios: Scalable Multi-Accelerator FPGA Architecture for Efficient Inference of Transformer-based Models.
.- A Hierarchical Adaptive Posit-CIM Architecture for Edge Transformer Inference.
.- Kernel Acceleration.
.- DOA: Dataflow Optimization for Attention on Multi-Core DSPs with Three-Level Memory Hierarchy.
.- ARROW: Adaptive Row Reorganization for Warp-Balanced SpMV.
.- HieraNTT: A Memory Hierarchy-Aware Data Access Architecture for Efficient Number Theoretic Transform on GPU.
.- Streaming Accelerators.
.- A Configurable Streaming Accelerator for LUT-Based Super-Resolution on FPGA.
.- TL-Sort: A Fully Pipelined Hardware Architecture for Sorting Without Run-Drain Stalls.
.- Hardware-Accelerated Streaming Graph Processing with Fast Refinement.
.- Secure Systems.
.- Pitfall: Uncovering and Exploiting the Store Forwarding Predictor on Intel CPUs.
.- InputSnatch: Stealing Input in LLM Services via Cache-Sharing Timing Side-Channel Attacks.
.- TArCS: Trusted and Attack-Resilient Clock Source with TEE and RDMA.
.- Data Efficiency.
.- CRAFT: Exploiting Precharge Cost Asymmetry for Adaptive DRAM Row Buffer Management.
.- Improving Compression Ratio of Lossy Compression for HPC via Modeling-Based Arithmetic Coding.
.- VecTEE: Compact TEE Metadata Caching for Efficient Secure Vector Computing.
.- Poster Papers.
.- AMXStencil: Boosting the Performance of Stencil Computations on AMX-Powered CPUs via Fusion.
.- S²M: Spatiotemporal-Cooperative Symbiotic Memory for Heterogeneous PQC Accel-eration.
.- WaferSim: A Simulation Infrastructure for LLM Service on Wafer-scale Chips.
.-OrbitGuard: Hierarchical Orbit-Aware Runtime for Spaceborne LLM Inference.
.- MV-BSD: Multi-Variable Speculation Diagrams for Compact Automated Logic De-sign.
.- RV-CIM: Energy–Delay Optimized Mapping and Architecture Co-Design for a RISC-V Multi-Core SoC with Configurable DCIM Cluster.
.- LLM-SYCL: Automated SYCL Generation from CUDA via Search-Driven LLM Translation.
.- Characterizing and Mitigating I/O Bottlenecks in LLM Inference on Disaggregated HPC Systems .
.- A Bottom-Up Approach for Die-to-Die Interconnect Optimization: From PHY Laten-cy Reduction to System-Level Design Implications.
.- EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints.
.- COMET: An FP32 Matrix Multiplication Accelerator by Extending INT8-based Ar-rays..
.- Work-in-Progress Papers.
.- RT-EdgeDetect: A Latency-Oriented Parallel YOLO Inference Framework for Re-source-Constrained Edge Devices.
.- GIST: Accelerating GCN Inference on 3D-Stacked PIM through Topology-Driven Mapping.
.- DMAgent: An Agentic Dynamic Disaggregated Memory Management Framework.
.- MixPIM: An ISA-Flexible and Mixed-Precision PIM Architecture for LLM Inference.




