Vibe Coding Discover

Use Cases

Multi-lora Adapter Serving

Published projects tagged with this use case.

1 project

vllm

★ 92K

vLLM is a high-throughput, memory-efficient LLM inference and serving engine built on PagedAttention, continuous batching, and optimized CUDA/ROCm kernels. It exposes an OpenAI-compatible server with quantization, LoRA, speculative decoding, and disaggregated serving across 200+ model architectures.

AI Frameworks | Python · inference-engine · llm-serving

View Project →