Featured Posts
Scaling Distributed GEMM on Cerebras Wafer-Scale Engine
Scaling Distributed GEMM on Cerebras Wafer-Scale Engine In Large Language Model (LLM), the fundamental operations of transformer architecture are attention and multi-layer perceptron computation, both of which are built on a massive amount of GEMM (General Matrix Multiply) and GEMV (General Matrix-Vector Multiplication). During inference, the decoding step (specifically GEMV) is memory-bandwidth bound due to the LLM autoregressive nature. (i.e. the GPU, such as NVIDIA’s accelerator, spends most of its time loading data into the compute unit for a relatively little computation, which computes one new token per step.) To address this issue, research has been attempting to transform this memory-bound workload to compute-bound by converting GEMV into GEMM operation through batching. However, memory movement is still relatively expensive. This observation of memory movement as the primary bottleneck, has encouraged companies like Cerebras to adopt a spatial dataflow architecture, in which compute units form a processing grid, enabling them to send data to one another at lower memory-transfer costs, thus addressing memory bound application from a hardware approach. ...
Optimizing Scientific Applications on HPC Systems
This blog offers an engineer’s perspective on optimizing the performance of scientific applications on HPC heterogeneous systems, drawing from international HPC competition experience and an internship in NSCC.