backburner
View on GitHubYour iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable
A llama.cpp-based local inference engine that splits Qwen3.8-27B workloads between an Apple Silicon Mac and a connected iPhone. It uses the phone for prompt processing and long-context attention, and exposes an OpenAI-compatible server.
Use Cases
Run a local LLM on Apple SiliconOffload inference work from a Mac to an iPhone over USB-CIncrease local model context capacityServe local models through an OpenAI-compatible APISpeed up decoding with speculative decoding
Built With
- Language
- Python
- Frameworks
- llama.cpp · Metal · Core ML
Tags
local LLM inference · llama.cpp · Apple Silicon · iPhone offload · split inference · long context · speculative decoding · Metal · SME2