llamAmpere
★ 158A llama.cpp fork optimized for Ampere GPUs, especially RTX 3090/3090 Ti. It adds tuned CUDA kernels, compressed KV-cache options, and multi-token prediction support for faster local inference and long contexts on consumer GPUs.
AI Frameworks | C++ · LLM inference · local LLM
View Project →