Vibe Coding Discover

AI Frameworks

backburner

View on GitHub

Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

★ 45141 forksPythonMITStayLameBro

A llama.cpp-based local inference engine that splits Qwen3.8-27B workloads between an Apple Silicon Mac and a connected iPhone. It uses the phone for prompt processing and long-context attention, and exposes an OpenAI-compatible server.

Use Cases

Run a local LLM on Apple SiliconOffload inference work from a Mac to an iPhone over USB-CIncrease local model context capacityServe local models through an OpenAI-compatible APISpeed up decoding with speculative decoding

Built With

Language
Python
Frameworks
llama.cpp · Metal · Core ML

Tags

local LLM inference · llama.cpp · Apple Silicon · iPhone offload · split inference · long context · speculative decoding · Metal · SME2