awesome-evals
★ 923An annotated collection of papers, articles, talks, tools, and benchmarks for building and evaluating AI agents. Includes a practical playbook with runnable examples for graders, trajectory evaluation, error analysis, and CI gating.
AI Tools | AI agent evaluation · LLM evaluation
View Project →Codex-QQ-Skin
★ 568A macOS and Windows utility that injects custom themes into the Codex/ChatGPT desktop app, with skin management, local token statistics, and an optional theme-generation skill. It uses local Chromium DevTools Protocol injection rather than modifying the official app bundle.
AI Tools | JavaScript · Codex · ChatGPT desktop
View Project →MinusPod
★ 458Self-hosted podcast server that transcribes episodes with Whisper, uses an LLM to identify ad segments, and serves edited audio through RSS feeds. Supports Anthropic, Ollama, OpenRouter, and OpenAI-compatible providers.
AI Tools | Python · podcasts · ad detection
View Project →agent-storyboard
★ 348A local storyboard workspace for Codex and Claude Code, with MCP tools for creating projects, managing shots, generating media, and returning results to the right shot. Includes voiceover alignment and HyperFrames or Remotion workflows.
AI Tools | JavaScript · AI video · storyboarding
View Project →antigravity-fixer
★ 130A Python CLI that diagnoses Google Antigravity and Gemini Code Assist eligibility and login errors, then cleans local credentials and cache and guides account re-authentication. Supports Windows, macOS, and Linux.
AI Tools | Python · Antigravity · Gemini Code Assist
View Project →awesome-ai-gateway
★ 119A bilingual catalog and comparison guide to 160+ AI gateways, with decision tools, reproducible cost benchmarks, security and compliance data, and an eight-chapter handbook. Helps developers choose and evaluate gateways for LLM routing and agent traffic.
AI Tools | HTML · AI gateways · LLM routing
View Project →guardana
★ 101Python security verification tool for AI artifacts, live model endpoints, MCP servers, and recorded agent traces. Runs offline artifact scans or bounded active probes, produces reproducible findings, and can gate CI with SARIF output.
AI Tools | HTML · AI security · LLM security
View Project →metrik
★ 100A cross-platform desktop app that tracks quotas, reset times, and token usage across locally installed AI coding agents. It parses local logs and official quota endpoints, with reports and no cloud service.
AI Tools | Rust · AI coding agents · token usage
View Project →