Vibe Coding Discover

Use Cases

Accelerate Dense Llm Matrix Multiplication

Published projects tagged with this use case.

1 project

DeepGEMM

★ 8.5K

A CUDA kernel library for high-performance LLM operations on NVIDIA SM90 and SM100 GPUs, including FP8/FP4/BF16 GEMMs, fused MoE, and attention-indexer scoring. Kernels are compiled at runtime through DeepJIT.

AI Frameworks | Cuda · CUDA kernels · GEMM

View Project →