halogen-flash-server
★ 790A specialized inference server for Qwen3.8-Flash-Next on AMD Strix Halo GPUs. It exposes an OpenAI-compatible API, supports long contexts and speculative decoding, and can load compatible GGUF files.
AI Frameworks | Shell · LLM inference · model serving
View Project →