Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather …
Tag:
Kernel
-
-
TECH
MoonMath AI Open-Sources a HIP Attention Kernel for AMD MI300X That Beats AITER v3 on Every Shape and Rounding Mode
by Techaiappby Techaiapp 9 minutes readMoonMath AI team has released a bf16 forward attention kernel for AMD’s MI300X GPU. It is written …