llama.cpp
View on GitHubLLM inference in C/C++
llama.cpp is a dependency-free C/C++ LLM and VLM inference engine built on ggml. It runs quantized GGUF models on CPU, CUDA, Metal, Vulkan and many other backends, and ships CLI, server, web UI and OpenAI-compatible API tools.
Use Cases
Run LLMs and VLMs locally on CPU, GPU, or NPUServe an OpenAI-compatible REST API from a local modelQuantize models to 1.5-8 bit for lower memory useConvert Hugging Face models to GGUF formatHybrid CPU+GPU inference for models larger than VRAMBuild lightweight C/C++ inference into apps and edge devicesRun models directly from Hugging Face with one CLI commandDistributed inference over RPC backend
Built With
- Language
- C++
- Frameworks
- ggml · CUDA · Metal · HIP/ROCm · Vulkan · SYCL · OpenVINO · CANN · WebGPU · OpenCL · MUSA · BLAS · PyTorch · Hugging Face Transformers
Tags
llm-inference · gguf · quantization · local-llm · cpu-inference · gpu-inference · cpp · openai-compatible-api · cuda · metal · vulkan · multimodal · edge-inference · no-dependencies · inference-engine · model-conversion