Vibe Coding Discover

Use Cases

Serve Models With An Openai-compatible Api

Published projects tagged with this use case.

1 project

llamAmpere

★ 158

A llama.cpp fork optimized for Ampere GPUs, especially RTX 3090/3090 Ti. It adds tuned CUDA kernels, compressed KV-cache options, and multi-token prediction support for faster local inference and long contexts on consumer GPUs.

AI Frameworks | C++ · LLM inference · local LLM

View Project →