Vibe Coding Discover

AI Tools

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

★ 22K1,823 forksPythonApache-2.0comet-ml

Apache-2.0, self-hostable LLM observability and evaluation platform from Comet. Provides deep agent/LLM tracing, datasets and experiments, LLM-as-a-judge metrics, CI/CD eval via PyTest, prompt playground, and production monitoring dashboards.

Use Cases

Trace multi-step agent and tool-call treesLog and inspect LLM calls in dev and prodRun offline experiments and datasets on LLM appsEvaluate RAG quality (answer relevance, context precision)Detect hallucinations and run moderation metricsLLM-as-a-judge scoring with feedback annotationsTest LLM pipelines in CI/CD via PyTestProduction monitoring dashboards for cost/latency/tokensOnline evaluation rules for live trafficPrompt playground for prompt/model experimentationOptimize prompts and agents with Agent OptimizerApply guardrails for safe/responsible AIQuery traces and run evals from coding agents via Opik MCP

Built With

Language
Python
Frameworks
LangChain · LlamaIndex · OpenAI SDK · Google ADK · AutoGen · Flowise AI · PyTest · MCP · TypeScript SDK · Pydantic

Tags

llm-observability · llm-evaluation · tracing · llmops · agent-tracing · prompt-management · llm-as-a-judge · rag-evaluation · monitoring · self-hosted · datasets · guardrails · dashboards · mcp-server · evaluation · apache-2.0