ollama
View on GitHubGet up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Ollama is a Go-based local LLM runtime that downloads, runs, and serves open models (Gemma, Qwen, DeepSeek, Llama) via a CLI and REST API on port 11434. It powers local inference for coding agents, chat UIs, and RAG apps, and supports custom Modelfiles and imports.
Use Cases
Run open-weight LLMs locally with a single CLI commandServe models over a local REST API on port 11434Chat with models like Gemma or Qwen from the terminalBack AI coding agents such as Claude Code, Codex, Copilot CLI, and OpenCode with local modelsBuild RAG or chatbot apps against a self-hosted LLM endpointImport and customize models with ModelfilesRun private/offline inference without sending data to cloud APIsTurn local models into a personal assistant via OpenClaw integrations
Built With
- Language
- Go
- Frameworks
- llama.cpp · MLX · GGUF · Docker · Python SDK · JavaScript SDK · CMake
Tags
llm · local-llm · inference · model-runner · llama.cpp · gguf · quantization · cli · rest-api · self-hosted · go · model-management · openai-compatible · embeddings · offline-ai · developer-tools