Model-Optimizer
View on GitHubA unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Python library for optimizing PyTorch, Hugging Face, and ONNX models with quantization, pruning, distillation, sparsity, NAS, and speculative decoding. Exports optimized checkpoints for deployment with TensorRT-LLM, TensorRT, vLLM, and SGLang.
Use Cases
Built With
- Language
- Python
- Frameworks
- PyTorch · ONNX · Hugging Face Transformers · Hugging Face Diffusers · TensorRT · TensorRT-LLM · vLLM · SGLang · Megatron-LM · Hugging Face Accelerate
Tags
model optimization · quantization · post-training quantization · quantization-aware training · pruning · distillation · speculative decoding · sparsity · neural architecture search · inference acceleration · model compression