Vibe Coding Discover

AI Frameworks

self-hosted-ai-stack

View on GitHub

Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper, WhisperLive, Kokoro, Embeddings, Docling and MCP Gateway. Local-first, private by default, with lightweight stacks, optional HTTPS and NVIDIA CUDA acceleration. Multi-arch: amd64, arm64.

★ 15630 forksShellMIThwdsl2

Docker Compose bundle that deploys a full local AI stack: Ollama for LLMs, LiteLLM gateway, AnythingLLM chat UI, embeddings/RAG, Whisper STT, Kokoro TTS, Docling parsing, and an MCP Gateway. Includes lightweight stack variants, optional HTTPS and CUDA GPU acceleration.

Use Cases

Deploy a private local ChatGPT-style chat UIDocument Q&A and RAG pipelines with embeddings and pgvectorSpeech-to-text to LLM to text-to-speech voice pipelinesAI coding assistant with MCP tool accessConvert PDFs/DOCX to structured Markdown via DoclingSelf-hosted OpenAI-compatible LLM gateway with key managementGPU-accelerated local inference with NVIDIA CUDALightweight modular stacks for chat, voice, RAG, or code

Built With

Language
Shell
Frameworks
Ollama · LiteLLM · AnythingLLM · MCP Gateway · Docling · Whisper · WhisperLive · Kokoro · Docker Compose · Caddy · PostgreSQL · pgvector

Tags

self-hosted · docker-compose · ollama · litellm · mcp · rag · local-llm · whisper · text-to-speech · embeddings · docling · cuda · privacy · multi-arch · openai-compatible · infrastructure

self-hosted-ai-stack — Vibe Coding Discover