DeepGEMM
View on GitHubDeepGEMM: clean and efficient BLAS kernel library on GPU
A CUDA kernel library for high-performance LLM operations on NVIDIA SM90 and SM100 GPUs, including FP8/FP4/BF16 GEMMs, fused MoE, and attention-indexer scoring. Kernels are compiled at runtime through DeepJIT.
Use Cases
Accelerate dense LLM matrix multiplicationRun grouped MoE GEMMsFuse MoE expert computation with communicationCompute MQA indexer logitsOptimize LLM training weight gradientsDevelop and benchmark NVIDIA GPU kernels
Built With
- Language
- Cuda
- Frameworks
- PyTorch · CUDA · CUTLASS · CuTe · DeepJIT
Tags
CUDA kernels · GEMM · LLM inference · FP8 · FP4 · BF16 · Mixture of Experts · attention · GPU optimization · runtime JIT