Vibe Coding Discover

Use Cases

Low-vram Lora/qlora Fine-tuning Of Large Models

Published projects tagged with this use case.

1 project

airllm

★ 35K

AirLLM is a Python inference engine that runs very large LLMs (70B up to 671B/2.8T) on a single 4GB GPU without quantization, distillation, or pruning by streaming model layers from disk. Optional 4/8-bit block-wise compression gives ~3x speedup, and it supports macOS/Apple Silicon plus low-VRAM training of 125B models

AI Frameworks | Jupyter Notebook · llm-inference · memory-optimization

View Project →