Soup
View on GitHubFine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
Soup is a Python CLI and web UI for configuring, fine-tuning, evaluating, and exporting LLMs. It supports methods such as LoRA and QLoRA, with layer streaming designed to train large models on low-VRAM GPUs.
Use Cases
Fine-tune LLMs from YAML configurationsTrain models on consumer GPUs with limited VRAMRun supervised fine-tuning and preference optimizationExport quantized models to GGUF for local inferenceEvaluate and serve fine-tuned models
Built With
- Language
- Python
- Frameworks
- PyTorch · Transformers · PEFT · TRL · MLX · Hugging Face · bitsandbytes · Accelerate · Ollama · llama.cpp
Tags
LLM fine-tuning · QLoRA · LoRA · SFT · DPO · local AI · low VRAM · layer streaming · GGUF · CLI · model export