CUDA Kernels, Chasing cuBLAS
A matrix-multiply kernel optimised rung by rung toward cuBLAS, ending with an honest account of why the last 15% is hard.
CUDA C++
Each one has a writeup: how it works, what I tried that didn’t, and where the numbers came from.
1 of 6 projects
A matrix-multiply kernel optimised rung by rung toward cuBLAS, ending with an honest account of why the last 15% is hard.