gufo
View on GitHubStrix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,628.52pp, 59.41tg single user, 157.22 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 70.56tg tok/s single user with DFlash2
A C++ inference engine optimized for AMD Strix Halo systems with ROCm. It serves supported language, audio, image, and video models through a local OpenAI-compatible API.
Use Cases
Run local LLM inference on AMD Strix Halo hardwareServe models through an OpenAI-compatible APIRun concurrent inference requestsPerform speech recognition and synthesisGenerate or edit imagesGenerate video and audio
Built With
- Language
- C++
- Frameworks
- ROCm · HIP · hipBLAS · hipBLASLt · rocBLAS · CMake · Nix
Tags
LLM inference · local AI · AMD Strix Halo · ROCm · GPU · multimodal · OpenAI-compatible API · speech recognition · text-to-speech · image generation · video generation