Vibe Coding Discover

AI Tools

Model-Optimizer

View on GitHub

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

★ 3.9K621 forksPythonApache-2.0NVIDIA

Python library for optimizing PyTorch, Hugging Face, and ONNX models with quantization, pruning, distillation, sparsity, NAS, and speculative decoding. Exports optimized checkpoints for deployment with TensorRT-LLM, TensorRT, vLLM, and SGLang.

Use Cases

Quantize models for faster inference and smaller checkpointsPrune and distill large language modelsApply quantization-aware training to recover model accuracyExport optimized models for TensorRT-LLM, vLLM, or SGLang deploymentOptimize diffusion models for lower-latency image and video generationBenchmark inference performance across optimized models

Built With

Language
Python
Frameworks
PyTorch · ONNX · Hugging Face Transformers · Hugging Face Diffusers · TensorRT · TensorRT-LLM · vLLM · SGLang · Megatron-LM · Hugging Face Accelerate

Tags

model optimization · quantization · post-training quantization · quantization-aware training · pruning · distillation · speculative decoding · sparsity · neural architecture search · inference acceleration · model compression

Model-Optimizer — Vibe Coding Discover