ongrid
View on GitHubAn ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.
Ongrid is a self-hosted ops AI agent in Go that coordinates specialist SRE/network/compute sub-agents to investigate alerts, correlate metrics, logs and traces, and pinpoint root cause, then act from Slack, Telegram, Lark or DingTalk. Includes RAG knowledge search, MCP server registry, skills catalog and approval-gated
Use Cases
Automated root cause analysis on alertsChatOps incident response from Slack/Telegram/Lark/DingTalkCorrelating metrics, logs and traces across hostsTopology blast-radius analysis before changesAudited reverse-tunnel shell into hostsRunbook and code search over a RAG knowledge vaultKubernetes cluster inspection and upgrade managementNetwork device discovery and SNMP pollingAlert-driven workflow automation with approval gatesRead-only host inspection via sandboxed tools
Built With
- Language
- Go
- Frameworks
- CloudWeGo Eino · Prometheus · Grafana · Loki · Tempo · OpenTelemetry · Qdrant · chi · GORM · Slack SDK · Lark SDK · DingTalk Stream SDK
Tags
aiops · devops · observability · incident-response · root-cause-analysis · chatops · kubernetes · mcp · rag · self-hosted · llm-agent · multi-agent · prometheus · grafana · opentelemetry · sre