Vibe Coding Discover

AI Frameworks

llamAmpere

View on GitHub

llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8, supporting full 262K ctx. Supports low memory (12gb) configs for 200k+ ctx

★ 15818 forksC++MITJakeATX

A llama.cpp fork optimized for Ampere GPUs, especially RTX 3090/3090 Ti. It adds tuned CUDA kernels, compressed KV-cache options, and multi-token prediction support for faster local inference and long contexts on consumer GPUs.

Use Cases

Run local LLM inference on Ampere GPUsServe models with an OpenAI-compatible APIGenerate from prompts with up to 262K context on a 24 GB GPURun long-context inference on 12 GB GPUsAccelerate coding, agentic, and RAG workloadsUse quantized KV caches to reduce VRAM usage

Built With

Language
C++
Frameworks
llama.cpp · GGML

Tags

LLM inference · local LLM · llama.cpp fork · CUDA · Ampere GPUs · RTX 3090 · long context · KV cache quantization · speculative decoding · multi-token prediction · GGUF · quantization