SIMDization of Small Tensor Multiplication Kernels for Wide SIMD Vector Processors

Rodrigues, Christopher; Phaosawasdi, Amarin; Wu, Peng

doi:10.1145/3178433.3178436

Cited by 3 publications

(1 citation statement)

References 12 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…For batched Cholesky factorization and Kalman filters, Lemaitre et al [34] propose a template system. Rodrigues et al [44] specify a small DSL for static tensor multiplications-even parallelizing error correction in 5G base stations [14] warrants a DSL. Likewise, there is a DSL for stencil operations, prime examples for memory-bound kernels and the importance of minimizing memory transformations through registers, in CUDA [56].…”

Section: Related Workmentioning

confidence: 99%

Flynn’s Reconciliation

Thuerck

Weber²,

Bifulco³

2021

ACM Trans. Archit. Code Optim.

View full text Add to dashboard Cite

A large portion of the recent performance increase in the High Performance Computing (HPC) and Machine Learning (ML) domains is fueled by accelerator cards. Many popular ML frameworks support accelerators by organizing computations as a computational graph over a set of highly optimized, batched general-purpose kernels. While this approach simplifies the kernels’ implementation for each individual accelerator, the increasing heterogeneity among accelerator architectures for HPC complicates the creation of portable and extensible libraries of such kernels. Therefore, using a generalization of the CUDA community’s warp register cache programming idiom, we propose a new programming idiom (CoRe) and a virtual architecture model (PIRCH), abstracting over SIMD and SIMT paradigms. We define and automate the mapping process from a single source to PIRCH’s intermediate representation and develop backends that issue code for three different architectures: Intel AVX512, NVIDIA GPUs, and NEC SX-Aurora. Code generated by our source-to-source compiler for batched kernels, borG, competes favorably with vendor-tuned libraries and is up to 2× faster than hand-tuned kernels across architectures.

show abstract

Section: Related Workmentioning

confidence: 99%