Vibe Coding Discover

AI Frameworks

AirLLM 70B inference with single 4GB GPU

★ 35K3,680 forksJupyter NotebookApache-2.0lyogavin

AirLLM is a Python inference engine that runs very large LLMs (70B up to 671B/2.8T) on a single 4GB GPU without quantization, distillation, or pruning by streaming model layers from disk. Optional 4/8-bit block-wise compression gives ~3x speedup, and it supports macOS/Apple Silicon plus low-VRAM training of 125B models

Use Cases

Run 70B LLM inference on a single 4GB GPUServe 671B DeepSeek-V3 on ~12GB VRAMInference on Apple Silicon MacsLow-VRAM LoRA/QLoRA fine-tuning of large modelsSpeed up inference 3x with block-wise quantizationLayer-wise model sharding to diskCPU inference of sharded modelsRun open-source flagship models on consumer hardware

Built With

Language
Jupyter Notebook
Frameworks
PyTorch · Hugging Face Transformers · bitsandbytes · PEFT · Accelerate · safetensors · MLX · flash-attn · compressed-tensors · einops · wandb · scikit-learn

Tags

llm-inference · memory-optimization · low-vram · quantization · model-compression · layer-streaming · gpu · macos · apple-silicon · bitsandbytes · lora-finetuning · fp8 · prefetching · pytorch · inference-engine