DeepGEMM
★ 8.5KA CUDA kernel library for high-performance LLM operations on NVIDIA SM90 and SM100 GPUs, including FP8/FP4/BF16 GEMMs, fused MoE, and attention-indexer scoring. Kernels are compiled at runtime through DeepJIT.
AI Frameworks | Cuda · CUDA kernels · GEMM
View Project →