| logic-pro-mcp | ★ 100 | MCP | Swift | Logic Pro · music production · DAW automation · desktop automation · MIDI · stateful control · fail-closed · live readback | A local MCP server that lets AI clients control Logic Pro through structured tools and inspect live project state. Mutations use explicit targets, confirmation checks, and readback to report confirmed, uncertain, or failed outcomes. | View Project → |
| tuweb.dev | ★ 100 | AI Coding | Astro | autonomous coding · AI code generation · Claude Code · human voting · automated deployment · rollback · prompt security | A community-driven website where visitors vote on feature ideas and Claude Code implements the winning proposal every 30 minutes. The project includes prompt moderation, diff restrictions, build checks, deployment verification, and automatic rollback. | View Project → |
| ScienceBuddy | ★ 100 | AI Agents | Python | scientific-agents · interactive-agent · recursive-self-improvement · agent-harness · llm-agents · reinforcement-learning · grpo · agent-evaluation · biomedicine · tool-use · research-automation · trajectory-inspection · verifier · python | Research code and preview for ScienceBuddy, an interactive scientific agent workspace plus a double-recursive self-improvement experiment that alternates harness refinement with SkyRL GRPO model training on frozen scientific tasks. | View Project → |
| End-to-End-Agentic-Ai-Automation-Lab | ★ 100 | AI Agents | Jupyter Notebook | agentic-ai · multi-agent · langgraph · autogen · mcp · rag · langchain · n8n · vllm · memory · human-in-the-loop · guardrails · fine-tuning · deployment · docker · browser-automation | A large hands-on lab of Jupyter notebooks and Python projects covering LangGraph and AutoGen agent systems, MCP servers, production RAG with reranking, mem0 memory, n8n automation, plus fine-tuning and vLLM deployment. Best for developers learning end-to-end agentic AI by example. | View Project → |
| winnow | ★ 100 | AI Coding | Python | claude-code · claude-code-plugin · context-management · hooks · token-reduction · context-pruning · llm-agents · relevance-judge · calibrated-model · mcp · python · shadow-mode | A Claude Code plugin that judges each large tool result with a calibrated System One model (Jev, or a Haiku adapter) and replaces unneeded blocks with a stub plus a recall key. Hooks, shadow mode, and MCP server included; local cache keeps the hidden text restorable. | View Project → |
| Canny | ★ 85 | AI Coding | TypeScript | coding-agent-guardrails · agent-hooks · evidence-ledger · done-gate · deterministic-checks · secret-detection · Claude Code · Codex CLI | A TypeScript supervision layer for Claude Code and Codex CLI. Hooks record agent actions and block completion when code changes lack a passing check; deterministic rules can also catch secrets and test removal. Jev-backed judgments are advisory. | View Project → |
| cyber-xiaowan | ★ 84 | Skills | JavaScript | skill-pack · claude-code · codex · prompt-pack · thinking-framework · cli · npx · installer · desktop-pet · privacy · progressive-disclosure · chinese · nodejs · agent-skill · content-editing · decision-support | Cyber Xiaowan is a public SKILL pack for Codex/Claude Code that acts as a lightweight structured-thinking operating system: it listens, clarifies, audits, architects, or suggests one small next step. Ships an npx installer, progressive-loading reference rules under references/, and an experimental local desktop-pet mod | View Project → |
| AMA-Bench | ★ 82 | AI Agents | Python | agent-memory · long-horizon · long-context · benchmark · evaluation · llm-as-judge · memory-retrieval · embeddings · bm25 · vllm · leaderboard · agent-trajectories · icml · python | AMA-Bench is an ICML 2026 evaluation framework for agentic memory: methods build memory from long agent trajectories, retrieve evidence, and answer QA scored by LLM-as-judge. Includes vLLM/API pipelines, cross-judge validation, and a HF leaderboard. | View Project → |
| blink | ★ 79 | AI Tools | TypeScript | code search · natural language · codebase navigation · CLI · developer tools · file discovery · TypeScript · Jev | Blink is a TypeScript CLI that uses Jev to search codebases with natural-language queries. It ranks likely files by walking directory paths and can run multiple walkers to explore recursively. | View Project → |
| jev-curate | ★ 74 | AI Tools | Rust | dataset curation · synthetic data · pretraining · fine-tuning · LLM evaluation · reasoning data · Parquet · JSONL · streaming · CLI · Rust · Python | A Rust CLI and Python-bindable pipeline that streams Parquet or JSONL datasets and uses TypeSafe AI's Jev model to score, filter, and curate examples. Includes presets for reasoning quality, anti-sycophancy, and code correctness. | View Project → |
| semdecide | ★ 73 | AI Tools | Python | cli · unix-pipelines · semantic-judgment · llm-classification · jsonl · ci · agent-safety · guardrails · exit-codes · python · shell-integration · typed-decisions · confidence-threshold · prompt-free | Python CLI that brings typed LLM judgments to Unix pipelines and CI: is/choose/score/filter commands return predicates, routes, scores, and stable exit codes from TypeSafe AI Jev, plus a guard recipe for agent action safety. | View Project → |
| MemoryArena | ★ 64 | AI Agents | Python | agent-memory · benchmark · multi-session · llm-agents · evaluation · memory-systems · long-context · tool-use · environments · research | MemoryArena is a research framework and benchmark for agent memory in interdependent multi-session agentic tasks. It wires pluggable memory backends (long-context, mem0, Letta, Mirix, GraphRAG, MemoRAG, BM25) into task agents and step-based environments for web shopping, travel, search, and formal reasoning. | View Project → |
| jev-mcp | ★ 55 | MCP | TypeScript | MCP server · coding workflow · typed judgments · patch review · tool routing · verification · stdio | A local stdio MCP server that exposes TypeSafe Jev judgments to Cursor, Codex, and other MCP clients. It supports coding-step routing, prepared-call selection, patch review, evidence checks, content screening, and candidate ranking; the host agent executes tools and edits files. | View Project → |
| curlora | ★ 53 | AI Tools | Jupyter Notebook | lora · peft · fine-tuning · continual-learning · catastrophic-forgetting · cur-decomposition · llm · parameter-efficient · matrix-decomposition · research-code · sft · quantization-planned · transformers · pytorch | Research code for CURLoRA, a LoRA variant that uses CUR matrix decomposition to fine-tune LLMs with fewer trainable parameters and less catastrophic forgetting. Includes PyTorch/Transformers implementations and notebooks reproducing experiments on Mistral 7B and GPT2-Large. | View Project → |
| HOMER | ★ 45 | AI Frameworks | Python | long-context · kv-cache · context-extension · llm-inference · llama-2 · memory-efficiency · attention · transformers · pytorch · research-implementation · hierarchical-merging · passkey-retrieval · perplexity-evaluation · training-free · flash-attention · iclr-2024 | Official ICLR 2024 implementation of HOMER, a training-free hierarchical KV-cache merging method that extends pre-trained LLM context limits (e.g. Llama-2) with lower memory. Ships patched LlamaForCausalLM, plus passkey-retrieval and PG19 perplexity scripts. | View Project → |
| OneVOneJev | ★ 38 | AI Agents | TypeScript | AI opponent · game agent · FPS · quickscope · real-time · WebSocket · spectator mode · heuristic fallback | A browser-based 1v1 quickscope FPS where players fight Jev, an AI opponent driven by the TypeSafe SDK. The server runs the match simulation and uses a deterministic heuristic fallback when the AI service is unavailable. | View Project → |
| agency-agents | ★ 29 | AI Agents | Shell | ai-agents · subagents · agent-personas · claude-code · cursor · codex · gemini-cli · prompt-engineering · multi-agent · agent-roster · developer-tools · copilot · windsurf · aider · markdown-prompts · workflow-automation | A MIT-licensed roster of 200+ markdown AI agent personas (engineering, design, marketing, sales, security, GIS, game-dev) installable as subagents into Claude Code, Cursor, Codex, Gemini CLI and other agentic tools via shell scripts or a desktop app. | View Project → |
| jev-trader | ★ 20 | AI Agents | Python | AI trading · market making · autonomous trading · risk management · model calibration · paper trading · deterministic state | A Python market-making system that uses Jev for selected market judgments while deterministic code handles state, quote policy, execution, and hard risk vetoes. Includes paper trading, fallback behavior, and decision calibration. | View Project → |
| brain-portal | ★ 19 | AI Tools | TypeScript | ai · knowledge-base · notes · rag · embeddings · mcp · self-hosted · vector-search · agents · nextjs · turso · sqlite-fts5 · entity-graph · openrouter · personal-crm · voice-capture | Self-hosted AI-native knowledge base for notes, tasks, projects and contacts. Every item is embedded and FTS5-indexed for semantic Q&A, and a 38-tool MCP server lets Claude Code, Cursor or other agents read and capture into your workspace. | View Project → |
| character-animation-skill | ★ 17 | Skills | Python | agent-skill · image-to-video · character-animation · sprite-sheet · frame-extraction · ffmpeg · python · imagegen · codex · game-assets · animation-pipeline · video-generation | Local Agent skill that turns a character reference image into a short action video, PNG animation frames, and a sprite sheet. Includes Python scripts for image-to-video generation, frame extraction, alignment, and spritesheet building. | View Project → |
| collabosm | ★ 15 | AI Tools | Python | LLM inference · model serving · OpenAI-compatible API · Google Colab · A100 · ExLlamaV3 · KV cache · Qwen | Scripts to deploy Qwen3.8-Flash-Next on a high-RAM Colab A100 and expose it through an authenticated OpenAI-compatible API. Includes runtime and model setup, session recovery, performance measurements, and KV-cache management. | View Project → |
| opencode-agents | ★ 12 | AI Coding | TypeScript | opencode · ai-agents · coding-agent · prompts · agent-config · multi-agent · subagents · plan-first · workflow · cli · typescript · code-review · automated-testing · evals · system-builder · installer | OpenCode agent framework: installable prompts, agents, subagents and commands for plan-first, approval-based coding workflows, plus an eval harness that runs agent test suites against Claude, GPT-4 and Grok. Built for the OpenCode CLI, not standalone. | View Project → |
| askgrokwallet | ★ 8 | AI Agents | JavaScript | agent-governance · agentic-commerce · ai-agents · human-in-the-loop · policy-engine · verifiable-receipts · wallet · ed25519 · x402 · erc-8196 · erc-8126 · mcp · approval-workflow · audit-log · base · smart-contracts | Rules-and-receipts layer for AI agents that spend money: plain-English policy compiles to allow/ask/deny, risky actions hit a human approval inbox, and every outcome produces an Ed25519-signed receipt anchored on Base. Contracts are unaudited; mainnet settlement not yet demonstrated. | View Project → |
| typesafe-ai-playground | ★ 5 | AI Tools | Rust | rust · cli · llm-api · phi-detection · code-review · line-importance · tui · text-classification · tree-sitter · syntax-highlighting · streaming · healthcare · interactive · scoring · terminal-ui | Rust CLI playground for TypeSafe's Jev (System One) model API. Commands score text for PHI, rate code-comment accuracy and usefulness, build line-importance heat maps, analyze typed tone live, and classify businesses or occupations. | View Project → |
| customer-service-agent | ★ 1 | AI Agents | Python | customer-service · langgraph · rag · booking · appointment-scheduling · ragflow · fastapi · stripe-payments · sse-streaming · voice · whisper-stt · tts · postgres · human-handoff · ticketing · single-tenant | ServiceEmma is a LangGraph customer-service agent for appointment businesses: it answers FAQs from RAG knowledge, books or cancels appointments, holds slots via Stripe Checkout, and escalates to support tickets. Ships with FastAPI backend, SSE chat widget, Postgres state, voice, and an owner dashboard. | View Project → |