self-hosted-ai-stack
View on GitHubDeploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper, WhisperLive, Kokoro, Embeddings, Docling and MCP Gateway. Local-first, private by default, with lightweight stacks, optional HTTPS and NVIDIA CUDA acceleration. Multi-arch: amd64, arm64.
Docker Compose bundle that deploys a full local AI stack: Ollama for LLMs, LiteLLM gateway, AnythingLLM chat UI, embeddings/RAG, Whisper STT, Kokoro TTS, Docling parsing, and an MCP Gateway. Includes lightweight stack variants, optional HTTPS and CUDA GPU acceleration.
Use Cases
Deploy a private local ChatGPT-style chat UIDocument Q&A and RAG pipelines with embeddings and pgvectorSpeech-to-text to LLM to text-to-speech voice pipelinesAI coding assistant with MCP tool accessConvert PDFs/DOCX to structured Markdown via DoclingSelf-hosted OpenAI-compatible LLM gateway with key managementGPU-accelerated local inference with NVIDIA CUDALightweight modular stacks for chat, voice, RAG, or code
Built With
- Language
- Shell
- Frameworks
- Ollama · LiteLLM · AnythingLLM · MCP Gateway · Docling · Whisper · WhisperLive · Kokoro · Docker Compose · Caddy · PostgreSQL · pgvector
Tags
self-hosted · docker-compose · ollama · litellm · mcp · rag · local-llm · whisper · text-to-speech · embeddings · docling · cuda · privacy · multi-arch · openai-compatible · infrastructure