lemonade
View on GitHubLemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Lemonade is a local AI server that runs optimized LLMs, speech, and image models on your own GPU/NPU, exposing OpenAI, Anthropic, and Ollama compatible APIs. Ships a CLI, model manager, and MCP server for connecting desktop apps and coding agents to private on-device inference.
Use Cases
Serve local LLMs to apps via OpenAI/Anthropic/Ollama APIsRun chat and coding models offline on GPU or NPULocal image generation with Stable DiffusionText-to-speech with KokoroReal-time speech transcription with WhisperEmbed local multimodal AI into desktop appsExpose local models to MCP-compatible agentsCloud offload routing to OpenAI/Fireworks/OpenRouterPoint coding assistants like Claude Code and Copilot at local modelsDownload and manage GGUF/FLM/ONNX models
Built With
- Language
- C++
- Frameworks
- llama.cpp · ONNX Runtime · ROCm · Vulkan · Metal · CUDA · Docker · CMake
Tags
local-llm · llm-inference · openai-api · onnxruntime · llamacpp · npu · rocm · vulkan · multimodal · gguf · text-to-speech · image-generation · mcp-server · amd · local-server · quantization