airllm
View on GitHubAirLLM 70B inference with single 4GB GPU
AirLLM is a Python inference engine that runs very large LLMs (70B up to 671B/2.8T) on a single 4GB GPU without quantization, distillation, or pruning by streaming model layers from disk. Optional 4/8-bit block-wise compression gives ~3x speedup, and it supports macOS/Apple Silicon plus low-VRAM training of 125B models
Use Cases
Run 70B LLM inference on a single 4GB GPUServe 671B DeepSeek-V3 on ~12GB VRAMInference on Apple Silicon MacsLow-VRAM LoRA/QLoRA fine-tuning of large modelsSpeed up inference 3x with block-wise quantizationLayer-wise model sharding to diskCPU inference of sharded modelsRun open-source flagship models on consumer hardware
Built With
- Language
- Jupyter Notebook
- Frameworks
- PyTorch · Hugging Face Transformers · bitsandbytes · PEFT · Accelerate · safetensors · MLX · flash-attn · compressed-tensors · einops · wandb · scikit-learn
Tags
llm-inference · memory-optimization · low-vram · quantization · model-compression · layer-streaming · gpu · macos · apple-silicon · bitsandbytes · lora-finetuning · fp8 · prefetching · pytorch · inference-engine