Fastest GPU kernels, written from scratch.

Matrix Multiplication

Matrix multiplication of square bf16 matrices, accumulated in fp32.

N=4096 Kernel: 763 TFLOPs cuBLAS: 716 TFLOPs N=8192 Kernel: 808 TFLOPs cuBLAS: 795 TFLOPs

Explanation in https://cudaforfun.substack.com/p/outperforming-cublas-on-h100-a-worklog

To run:

make matmul && out/matmul

Example kernels are in examples/matmul/ and orchestration is in matmul.cu

Sum reduction

We compute sum of 2^30 elements.

To run:

make sum && out/sum

Kernel: 3240.11 GB/s cub Library: 3193 GB/s

Example kernels are in sum.cu

Name		Name	Last commit message	Last commit date
Latest commit History 16 Commits
examples/matmul		examples/matmul
.gitignore		.gitignore
LICENSE		LICENSE
Makefile		Makefile
README.md		README.md
logs.txt		logs.txt
matmul.cu		matmul.cu
sum.cu		sum.cu

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

Fastest GPU kernels, written from scratch.

Matrix Multiplication

To run:

Sum reduction

To run:

About

Uh oh!

Releases

Packages

Uh oh!

Contributors 3

Languages

License

pranjalssh/fast.cu

Folders and files

Latest commit

History

Repository files navigation

Fastest GPU kernels, written from scratch.

Matrix Multiplication

To run:

Sum reduction

To run:

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors 3

Languages

Packages