Vibe Coding Discover

Use Cases

Quantize Models For Faster Inference And Smaller Checkpoints

Published projects tagged with this use case.

1 project

Model-Optimizer

★ 3.9K

Python library for optimizing PyTorch, Hugging Face, and ONNX models with quantization, pruning, distillation, sparsity, NAS, and speculative decoding. Exports optimized checkpoints for deployment with TensorRT-LLM, TensorRT, vLLM, and SGLang.

AI Tools | Python · model optimization · quantization

View Project →