Vibe Coding Discover

AI Frameworks

DeepGEMM: clean and efficient BLAS kernel library on GPU

★ 8.5K1,359 forksCudaMITdeepseek-ai

A CUDA kernel library for high-performance LLM operations on NVIDIA SM90 and SM100 GPUs, including FP8/FP4/BF16 GEMMs, fused MoE, and attention-indexer scoring. Kernels are compiled at runtime through DeepJIT.

Use Cases

Accelerate dense LLM matrix multiplicationRun grouped MoE GEMMsFuse MoE expert computation with communicationCompute MQA indexer logitsOptimize LLM training weight gradientsDevelop and benchmark NVIDIA GPU kernels

Built With

Language
Cuda
Frameworks
PyTorch · CUDA · CUTLASS · CuTe · DeepJIT

Tags

CUDA kernels · GEMM · LLM inference · FP8 · FP4 · BF16 · Mixture of Experts · attention · GPU optimization · runtime JIT