Strata
View on GitHubQwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
Strata runs Qwen3.8-Flash-Next on consumer NVIDIA or AMD hardware, using GPU, RAM, and SSD resources to fit the large model. It provides a browser chat UI and local OpenAI-compatible, Anthropic-compatible, and MCP interfaces.
Use Cases
Run a large language model locally on consumer hardwareServe a local model through OpenAI-compatible APIsConnect coding agents to a local modelChat with a local model using a browser UIAnalyze images with a local model
Built With
- Language
- C++
- Frameworks
- llama.cpp · ggml · MCP
Tags
local inference · LLM serving · consumer hardware · GPU acceleration · OpenAI-compatible API · Anthropic API · multimodal · MCP server · Windows · Linux