halogen-flash-server
View on GitHubThe fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
A specialized inference server for Qwen3.8-Flash-Next on AMD Strix Halo GPUs. It exposes an OpenAI-compatible API, supports long contexts and speculative decoding, and can load compatible GGUF files.
Use Cases
Serve Qwen3.8-Flash-Next locally on AMD Strix HaloProvide an OpenAI-compatible endpoint for applications and agentsRun long-context inferenceUse speculative decoding to accelerate generationBenchmark and evaluate model checkpoints
Built With
- Language
- Shell
- Frameworks
- ROCm · HIP · llama.cpp GGUF
Tags
LLM inference · model serving · local LLM · AMD GPU · ROCm · Strix Halo · Qwen · OpenAI-compatible API · speculative decoding · long context · GGUF