tensorflow
★ 200KAn end-to-end machine-learning framework with Python and C++ APIs for building, training, and deploying models across CPUs, GPUs, and other devices.
AI Frameworks | C++ · machine-learning · deep-learning
View Project →Category
Core frameworks for building LLM applications.
An end-to-end machine-learning framework with Python and C++ APIs for building, training, and deploying models across CPUs, GPUs, and other devices.
AI Frameworks | C++ · machine-learning · deep-learning
View Project →Ollama is a Go-based local LLM runtime that downloads, runs, and serves open models (Gemma, Qwen, DeepSeek, Llama) via a CLI and REST API on port 11434. It powers local inference for coding agents, chat UIs, and RAG apps, and supports custom Modelfiles and imports.
AI Frameworks | Go · llm · local-llm
View Project →Hugging Face Transformers is the model-definition framework for state-of-the-art text, vision, audio, video and multimodal models. It offers a unified Pipeline/Trainer API over 1M+ Hub checkpoints and feeds model definitions to vLLM, llama.cpp and training stacks.
AI Frameworks | Python · transformers · model-definition
View Project →Dify is an open-source LLM app platform combining a visual agentic workflow builder, RAG pipeline, prompt IDE, agent runtime, and LLMOps. It supports hundreds of models, MCP servers, and tools, and self-hosts via Docker with APIs for integration.
AI Frameworks | TypeScript · llm · agents
View Project →Langflow is an open-source visual platform for building and deploying LLM agents and workflows. It combines a drag-and-drop React Flow canvas with Python component customization, then ships each flow as a REST API, an MCP server, or exportable JSON.
AI Frameworks | Python · visual-builder · low-code
View Project →LangChain is the Python/JS framework for building LLM apps and agents: a standard interface for chat models, embeddings, vector stores, tools, and retrievers, plus chains, structured output, and streaming. Paired with LangGraph for controllable agent workflows.
AI Frameworks | Python · agents · llm
View Project →llama.cpp is a dependency-free C/C++ LLM and VLM inference engine built on ggml. It runs quantized GGUF models on CPU, CUDA, Metal, Vulkan and many other backends, and ships CLI, server, web UI and OpenAI-compatible API tools.
AI Frameworks | C++ · llm-inference · gguf
View Project →High-performance C/C++ inference for local LLMs across CPUs, GPUs, and devices.
AI Frameworks | C++ · llama.cpp · Inference
View Project →vLLM is a high-throughput, memory-efficient LLM inference and serving engine built on PagedAttention, continuous batching, and optimized CUDA/ROCm kernels. It exposes an OpenAI-compatible server with quantization, LoRA, speculative decoding, and disaggregated serving across 200+ model architectures.
AI Frameworks | Python · inference-engine · llm-serving
View Project →Unsloth is a Python training and inference stack plus desktop/web UI for running, fine-tuning and RL-training LLMs and diffusion models 2x faster with less VRAM. Supports LoRA/QLoRA/full fine-tuning, GGUF/MLX export, local OpenAI-compatible serving, RAG and agent hooks. Apache-2.0.
AI Frameworks | Python · fine-tuning · llm
View Project →LabML's annotated deep learning library: 60+ readable PyTorch implementations of papers with side-by-side notes, covering transformers (GPT, ViT, XL), LoRA, diffusion/Stable Diffusion, GANs, optimizers, and RL. Best as a learning and reference resource rather than a production framework.
AI Frameworks | Python · transformers · pytorch
View Project →Minimal, readable ~300-line PyTorch repo for training and finetuning GPT-2-scale models; reproduces GPT-2 124M on OpenWebText. Author marks it deprecated in favor of the newer nanochat, but it remains a compact hacking/education base.
AI Frameworks | Python · gpt · llm-training
View Project →LiteLLM is an open-source AI gateway and Python SDK that calls 100+ LLM providers (OpenAI, Anthropic, Bedrock, Vertex, vLLM, Ollama) in OpenAI format. The proxy adds virtual keys, spend tracking, guardrails, load balancing, logging, plus MCP tool and A2A agent routing.
AI Frameworks | Python · ai-gateway · llm-gateway
View Project →LlamaIndex is a Python framework for building LLM apps: data connectors, indices, retrievers and query engines for RAG, plus agent/workflow abstractions. 300+ integrations let you swap LLMs, embeddings and vector stores; the team now also pushes the hosted LlamaParse document platform.
AI Frameworks | Python · rag · llm
View Project →LocalAI is a self-hosted, OpenAI/Anthropic-compatible inference engine that runs LLMs, vision, voice, image and video models on CPU or GPU. Backends like llama.cpp, vLLM and whisper.cpp are pulled on demand, and it also ships agents, RAG and MCP support.
AI Frameworks | Go · llm-inference · local-inference
View Project →Agno is a Python framework plus AgentOS runtime for building, serving, and managing agent platforms. It adds REST/SSE endpoints, storage, memory, RAG knowledge, 100+ toolkits, human approval, scheduling, RBAC and OpenTelemetry observability, with Docker and cloud deploy templates.
AI Frameworks | Python · agents · multi-agent
View Project →LangGraph is a low-level Python orchestration framework for building stateful, long-running agents as graphs. It adds durable execution, checkpointing, human-in-the-loop interrupts, and memory, with optional LangSmith tracing and deployment.
AI Frameworks | Python · agents · multi-agent
View Project →DSPy is a Python framework for programming rather than prompting LLMs: you write declarative modules and signatures, then compilers/optimizers (e.g. GEPA, MIPRO) tune prompts and weights against your metric. Useful for building and optimizing RAG pipelines, classifiers, and agent loops.
AI Frameworks | Python · dspy · prompt-optimization
View Project →Colibrì is a pure-C, zero-dependency inference engine that runs frontier MoE models (GLM-5.x 744B, Kimi K3 2.8T, DeepSeek V4, Qwen3.x) on consumer hardware by streaming disk-resident experts into a unified VRAM/RAM/NVMe tier hierarchy. It offers chat/serve/web front ends plus CPU, CUDA, Metal and Vulkan backends.
AI Frameworks | C · inference-engine · moe
View Project →CopilotKit is a TypeScript SDK for building agent-native apps: chat UI, generative UI, shared state, and human-in-the-loop across React, Angular, Vue, React Native, Slack, and Teams. It also maintains the AG-UI protocol connecting agent frameworks to frontends.
AI Frameworks | TypeScript · agent-native · generative-ui
View Project →SGLang is a high-performance serving framework for LLMs and multimodal models, offering RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, and quantization over an OpenAI-compatible API across NVIDIA, AMD, TPU, and CPU hardware.
AI Frameworks | Python · llm-serving · inference-engine
View Project →AirLLM is a Python inference engine that runs very large LLMs (70B up to 671B/2.8T) on a single 4GB GPU without quantization, distillation, or pruning by streaming model layers from disk. Optional 4/8-bit block-wise compression gives ~3x speedup, and it supports macOS/Apple Silicon plus low-VRAM training of 125B models
AI Frameworks | Jupyter Notebook · llm-inference · memory-optimization
View Project →Fully local, open-source ElevenLabs alternative for voice cloning, voice design, dubbing, dictation, transcription and audiobooks in 646 languages. Ships a desktop app (Electron/Tauri), model manager, local API and an MCP server for agent integration.
AI Frameworks | Python · voice-cloning · text-to-speech
View Project →AgentScope 2.0 is Alibaba Tongyi's Python framework for building and running production agents: ReAct agents, toolkits over Python/MCP/skills, multi-agent teams, realtime voice, sandboxed execution, HITL, memory, plus a FastAPI agent service with web UI and IM channels.
AI Frameworks | Python · multi-agent · agent-framework
View Project →Modular's open-source platform hosting MAX (an inference framework with an OpenAI-compatible server and accelerator kernels) and the Mojo language, including its compiler, standard library, and Python model pipelines.
AI Frameworks | Mojo · max · llm-inference
View Project →Microsoft's model-agnostic SDK for building AI agents and multi-agent systems in Python, .NET, and Java. Ships plugins/tool calling, memory and vector connectors, planners, and MCP support; now succeeded by Microsoft Agent Framework.
AI Frameworks | C# · ai-agents · multi-agent
View Project →Mastra is a TypeScript framework for building AI agents and LLM apps: agents, graph workflows, RAG, memory, MCP servers, evals, and observability, with model routing across 40+ providers and React/Next.js/Node integration.
AI Frameworks | TypeScript · ai-agents · workflows
View Project →Vercel AI SDK is a provider-agnostic TypeScript toolkit for building AI apps and agents: unified model APIs, streaming text, structured output, tool loops, and framework-agnostic UI hooks for React, Next.js, Svelte and Vue.
AI Frameworks | TypeScript · ai-sdk · llm
View Project →Haystack is deepset's open-source Python framework for building production LLM apps: modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Model- and vendor-agnostic, with built-in components for RAG, semantic search, evaluation, and async streaming.
AI Frameworks | Python · llm-orchestration · rag
View Project →Python library for fast, calibrated typed decisions over text, including classification, scoring, and yes/no probabilities in a single model pass. It routes requests to English or multilingual checkpoints and supports HTTP, MCP, and LangChain integrations.
AI Frameworks | Python · decision-model · classification
View Project →oMLX is an MLX-based LLM inference server for Apple Silicon with continuous batching, tiered hot-RAM/cold-SSD KV cache, multi-model serving, and VLMs/embeddings/rerankers behind an OpenAI-compatible API, managed from a macOS menu bar app and web dashboard.
AI Frameworks | Python · llm-inference · apple-silicon
View Project →Cross-platform ML inference and training accelerator. Runs models from PyTorch, TensorFlow, and classical ML libraries via ONNX, applying graph optimizations and hardware acceleration on CPU, GPU, or NPU. Widely used runtime for deploying fast, low-cost model inference.
AI Frameworks | C++ · onnx · inference-engine
View Project →Hugging Face PEFT implements parameter-efficient fine-tuning methods (LoRA, QLoRA, IA3, prompt tuning, adapters and more) on top of Transformers, Diffusers, Accelerate and TRL. It lets you train and serve large models with a fraction of the GPU memory and storage, with only small adapter checkpoints.
AI Frameworks | Python · peft · lora
View Project →Pydantic AI is a Python AI agent framework from the Pydantic team: a typed agent loop with validated structured outputs, tools, dependency injection, MCP, multi-agent support, realtime voice, image generation, embeddings and evals, with provider-agnostic model switching.
AI Frameworks | Python · agent-framework · llm
View Project →Vercel Labs' Generative UI framework: an LLM emits JSON specs constrained to a Zod-typed component catalog, which json-render streams and renders safely across React, Vue, Svelte, Solid, React Native, Next.js, Remotion, PDF, email, 3D and terminal targets.
AI Frameworks | TypeScript · generative-ui · json-spec
View Project →MIT-licensed AI pipeline engine with a multithreaded C++ runtime and 50+ Python-extensible nodes. Compose LLM, RAG and multi-agent workflows as portable JSON in VS Code or via Python/TypeScript/MCP SDKs, backed by 13+ model providers and 9 vector DBs.
AI Frameworks | Python · ai-pipeline · llm-workflows
View Project →Structured LLM outputs validated by Pydantic, with retries and streaming.
AI Frameworks | Python · Structured Output · Pydantic
View Project →Microsoft's open multi-language SDK for building, orchestrating and hosting production AI agents and multi-agent workflows in Python and .NET. Ships graph workflows, middleware, OpenTelemetry observability, declarative YAML agents and Foundry hosting.
AI Frameworks | Python · agents · multi-agent
View Project →Semantica is a Python graph-native context and knowledge-graph layer for AI systems: ingest data, build context graphs, and run deterministic graph reasoning with W3C PROV-O provenance. Adds graph RAG, decision audit trails, ontology governance, and MCP/LangChain/CrewAI integrations on top of your existing LLM stack.
AI Frameworks | Python · knowledge-graph · context-graph
View Project →Eino is ByteDance's Go LLM application framework, modeled on LangChain and Google ADK. It offers component abstractions (ChatModel, Tool, Retriever), a graph/compose orchestration layer with automatic streaming, an agent development kit with ReAct and multi-agent patterns, and interrupt/resume for human-in-the-loop.
AI Frameworks | Go · golang · llm-framework
View Project →LangChain4j is an idiomatic Java library for building LLM-powered JVM apps, with a unified API over 20+ model providers and 30+ embedding stores. It covers tool calling (incl. MCP), agents, RAG pipelines, chat memory and prompt templating, with Spring Boot and Quarkus integrations.
AI Frameworks | Java · jvm · llm
View Project →Needle is an 8-29 MB 2-bit foundation model and Python toolkit for on-device tool calling, structured JSON extraction and text embeddings. It ships inference, LoRA fine-tuning and per-platform builds for phones, wearables, robots, cars and microcontrollers.
AI Frameworks | Python · on-device-ai · edge-ai
View Project →LMCache is a vendor-neutral KV cache management layer for LLM inference. It stores and reuses KV cache across CPU RAM, disk and remote backends, cutting TTFT and boosting throughput for long-context, multi-turn agentic and RAG workloads on engines like vLLM.
AI Frameworks | Python · kv-cache · llm-inference
View Project →OpenRLHF is a Ray + vLLM distributed RLHF framework with an agent-based execution pipeline, supporting PPO, GRPO, REINFORCE++, DAPO, async RL, and VLM/multi-turn agent training for LLMs up to 70B+ parameters.
AI Frameworks | Python · RLHF · reinforcement-learning
View Project →Google's code-first Go framework for building, orchestrating, evaluating, and deploying AI agents. It supports multi-agent workflows and tool integrations, is optimized for Gemini, and can use other model providers.
AI Frameworks | Go · AI agents · multi-agent systems
View Project →Python framework/SDK for building agents on the Model Context Protocol. Fully implements MCP lifecycle (tools, resources, prompts, OAuth, sampling) and composes Anthropic's effective agent patterns, with optional Temporal-backed durable execution and cloud deployment.
AI Frameworks | Python · mcp · ai-agents
View Project →A portable C99 inference engine that runs the 2.78T-parameter Kimi K3 MoE model on one CPU with 8.24 GB peak RAM, streaming a 1.56 TB checkpoint from disk with no BLAS, framework, GPU or runtime dependencies.
AI Frameworks | C · llm-inference · cpu-inference
View Project →Kev trains LoRA adapters plus a pointer readout head on Qwen3 (0.6B/4B/8B) to answer typed yes/no, choice, and score questions over a document in one causal prefill pass, returning calibrated probabilities rather than text. Ships a FastAPI /v1/systemone server, frozen eval suites, and a Next.js playground.
AI Frameworks | Python · decision-model · calibration
View Project →MLX LM is a Python package for running and fine-tuning LLMs on Apple silicon via MLX. It offers CLI and Python APIs for generation, chat, LoRA/full fine-tuning, quantization, GGUF conversion, prompt caching, and an OpenAI-compatible server.
AI Frameworks | Python · llm-inference · apple-silicon
View Project →DeepSpec is a full-stack Python codebase for training and evaluating speculative-decoding draft models (DSpark, DFlash, Eagle3) against targets like Qwen3 and Gemma. It covers data prep, 8-GPU training, and benchmark evaluation, plus released checkpoints.
AI Frameworks | Python · speculative-decoding · llm-inference
View Project →Shimmy is a Rust inference server for local GGUF language models, with WebGPU acceleration through Airframe and OpenAI-compatible chat, completion, and streaming APIs. It runs as a single binary without Python or llama.cpp.
AI Frameworks | Rust · LLM inference · local AI
View Project →Native Apple Silicon MLX inference runtime for Laya typed-decision models. Returns choice probabilities, rubric scores and P(true) locally in ~7-14 ms per short question, with no PyTorch, tokenizer decoding or cloud API at runtime.
AI Frameworks | Python · mlx · apple-silicon
View Project →Higgsfield is an open-source GPU orchestration and ML training framework for multi-node training of billion-to-trillion parameter LLMs. It wraps PyTorch FSDP and DeepSpeed ZeRO-3 with node allocation, experiment queuing, monitoring, and GitHub Actions-driven deployment across cloud nodes.
AI Frameworks | Jupyter Notebook · distributed-training · gpu-orchestration
View Project →Lemonade is a local AI server that runs optimized LLMs, speech, and image models on your own GPU/NPU, exposing OpenAI, Anthropic, and Ollama compatible APIs. Ships a CLI, model manager, and MCP server for connecting desktop apps and coding agents to private on-device inference.
AI Frameworks | C++ · local-llm · llm-inference
View Project →Java/Spring Boot enterprise AI platform on Langchain4j: multi-provider LLM management, local RAG with Milvus/Weaviate/Qdrant, MCP tool and Skill integration, visual workflow orchestration, and Supervisor-mode multi-agent coordination with admin and user frontends.
AI Frameworks | Java · multi-agent · rag
View Project →RubyLLM is a Ruby-native AI framework giving one consistent API for 18+ providers: chat, streaming, embeddings, RAG, tools, agents, structured output, images, audio, video, and OCR, with first-class Rails integration and cost tracking.
AI Frameworks | Ruby · rails · llm-framework
View Project →Python framework and CLI for building, running and evaluating LLM agents and workflows, with first-class MCP (client/server, sampling, elicitations), Agent Skills, ACP and A2A support. Includes a TUI coding agent plus declarative agent/workflow definitions and broad model provider coverage.
AI Frameworks | Python · agents · mcp
View Project →LazyLLM is a Python low-code framework for assembling multi-agent LLM apps from modular pipelines (pipeline, parallel, switch, loop). It bundles RAG, tool-calling agents, one-click deployment, and unified online/local model fine-tuning and inference via vLLM, LightLLM or LMDeploy.
AI Frameworks | Python · multi-agent · llm
View Project →Stability AI's toolkit for training and running generative audio models: latent diffusion, autoencoders, and LMs with text/audio conditioning. Includes training scripts, fine-tuning/LoRA support, and a Gradio UI for Stable Audio Open.
AI Frameworks | Python · audio-generation · text-to-audio
View Project →GuppyLM is a ~9M parameter vanilla transformer trained from scratch to chat as a fish persona. It includes data generation, BPE tokenizer training, training loop, inference, ONNX/WASM browser demo, and Colab notebooks, making it a compact reference for building your own tiny LLM.
AI Frameworks | Python · llm · tiny-llm
View Project →Self-hosted inference server and production cluster for agent workloads. Serves embedding, retrieval, reranking, generation, OCR, extraction, and multimodal models through an OpenAI-compatible API, with on-demand model loading and integrations for popular AI frameworks and vector stores.
AI Frameworks | Python · inference server · model serving
View Project →Neo.mjs is a multi-threaded JavaScript application engine (worker-based runtime, JSON-first UI, zero-build ES modules), framed as the runtime 'body' hosting an AI agent swarm; the Agent OS, MCP servers and GraphRAG live mostly in the sibling neo-agent-brain repo.
AI Frameworks | JavaScript · frontend-framework · multi-threaded
View Project →Microsoft GenAIScript is a JavaScript/TypeScript framework for writing LLM prompts as code, orchestrating models, tools, MCP servers and agents. It includes built-in RAG vector search, structured output schemas, evals, and a VS Code extension plus CLI. Note: the repository is marked DEPRECATED.
AI Frameworks | TypeScript · prompt-as-code · llm-orchestration
View Project →A self-hosted framework for building industry news sites. It collects from configurable sources, uses LLMs to filter and score articles, writes summaries, clusters related coverage into events, and publishes ranked topics and briefings.
AI Frameworks | TypeScript · AI news aggregation · LLM content selection
View Project →MTPLX is an Apple Silicon LLM inference engine and Mac app that uses Qwen's native multi-token prediction heads for exact speculative decoding (1.6x-2.24x faster decode at any temperature). It serves local models over OpenAI- and Anthropic-compatible APIs for coding agents and chat.
AI Frameworks | Python · apple-silicon · mlx
View Project →NanoJev is a 0.6B Qwen3-based replica of Jev: a parallel decision model that returns probability distributions over supplied candidate actions with zero output-token decoding. Ships SFT training configs, an 18.7K-question dataset, an HTTP decision service and ViZDoom/Maze/Snake benchmarks.
AI Frameworks | Python · decision-model · parallel-decoding
View Project →Rust framework for building AI agents with function calling and serverless LLM tools. It provides QUIC-based tool routing, an OpenAI-compatible API, and infrastructure for geo-distributed inference.
AI Frameworks | Rust · AI agents · function calling
View Project →A PHP framework for building generative AI applications, with support for multiple LLM providers, embeddings, vector stores, agents, and question answering. Integrates with Laravel and Symfony.
AI Frameworks | PHP · generative AI · LLM
View Project →Qwen's 7B text-to-image and image-editing diffusion model (32-layer single-stream DiT) with native RGBA transparency, up to 10 reference images, mask/local edits, and 2K output. Ships Diffusers, ComfyUI, vLLM-Omni and SGLang integrations plus prompt-rewriting checkpoints.
AI Frameworks | Python · text-to-image · image-editing
View Project →vLLM-style inference server for Apple Silicon built on MLX, exposing OpenAI /v1/* and Anthropic /v1/messages from one process. Adds continuous batching, paged/prefix KV cache, structured output, MCP tool calling, and multimodal text, vision, audio, embeddings and rerank support.
AI Frameworks | Python · llm-inference · mlx
View Project →Python framework for building LLM applications with agent, tool-use, and multi-agent orchestration primitives. Supports local and cloud model providers, multimodal workflows, MCP, and knowledge graph pipelines.
AI Frameworks | Python · LLM · multimodal
View Project →Python framework and CLI harness for building, debugging, and deploying tool-using AI agents. Includes browser, shell, email, and file integrations, reusable skills, approval controls, and support for hosting agents that other agents can call.
AI Frameworks | Python · AI agents · multi-agent
View Project →Java enterprise server plus Vue admin console for Xiaozhi ESP32 voice hardware. Multi-LLM (OpenAI/ZhiPu/Ollama/Dify/Coze), local and cloud STT/TTS with voice cloning, WebSocket/MQTT realtime audio, MCP tools, RAG, OTA and device monitoring.
AI Frameworks | Java · esp32 · voice-assistant
View Project →Python library for building environments to evaluate and train LLMs and agents. It provides benchmark environments, agent harnesses, verifiers, rollout collection, and scalable execution for evaluation and reinforcement-learning workflows.
AI Frameworks | Python · agent-evaluation · LLM-evaluation
View Project →Python starter for training a small one-pass scorer that turns text plus a changing list of options into one probability per option, an independent alternative to TypeSafe's Jev. Ships train/eval/predict CLIs, byte and frozen Hugging Face encoder paths, and Doom/chess vision-scoring examples.
AI Frameworks | Python · one-pass scorer · system-one model
View Project →C#/.NET port of LangChain offering composable chains, prompt templates, document loaders, embeddings and vector stores for building LLM and RAG applications, closely mirroring the original Python abstractions.
AI Frameworks | C# · langchain · csharp
View Project →CLM serves a contrastively trained model that scores candidate actions against a state, with an API for typed decisions and ranking. Use it for agent action selection, tool routing, or verifying and ranking generated solutions.
AI Frameworks | Python · contrastive learning · state-action scoring
View Project →Official Python SDK for LM Studio: connect to a local LM Studio server to run chat, completion, tool-use, and schema-constrained inference with locally hosted LLMs, with sync and async clients plus model load/unload management.
AI Frameworks | Python · python-sdk · llm
View Project →Python framework for declaring and running typed, composable AI methods in .mthds files. It orchestrates multi-step LLM pipelines, model routing, document extraction, and structured outputs.
AI Frameworks | Python · AI workflows · DSL
View Project →A Rust framework for building LLM agents, typed task graphs, and streaming RAG pipelines. It includes indexing and query components, tool and MCP integrations, and connectors for LLM providers and vector stores.
AI Frameworks | Rust · LLM agents · RAG
View Project →Zero-runtime-dependency TypeScript SDK for production AI agents: durable resumable runs, long-term memory, hybrid RAG, MCP tool calling, approvals and swarm orchestration over one streaming API for Claude, GPT, Gemini, Grok, Mistral and DeepSeek on Node, Bun, Deno and edge.
AI Frameworks | TypeScript · agent-framework · durable-execution
View Project →.NET SDK for building AI agents and workflows with 30+ provider connectors, MCP and A2A support, vector DB integrations, multimodal IO, and a graph-based agent orchestration API. Works with local runtimes like vLLM, Ollama and LocalAI.
AI Frameworks | C# · dotnet · csharp
View Project →Swift-native agent runtime for building AI agents with type-safe @Tool macros, Apple Foundation Models on-device inference, composable sequential/parallel/routed workflows, memory, guardrails, streaming, MCP bridging, and OpenTelemetry tracing. Ships as a Swift Package for iOS, macOS, and Linux.
AI Frameworks | Swift · agents · multi-agent
View Project →Recipe and scripts for online iterative RLHF: SFT, reward modeling, vLLM response generation, reward annotation, and iterative DPO training loops. Reproduces LLaMA3-8B alignment comparable to Llama3-8B-Instruct using only open-source data.
AI Frameworks | Python · rlhf · dpo
View Project →SMG is a Rust LLM gateway that unifies OpenAI/Anthropic/Gemini and self-hosted engines (vLLM, SGLang, TensorRT-LLM, TokenSpeed, MLX) behind one API. It adds KV-cache-aware routing, gRPC pipelines, MCP tooling, multi-tenancy and observability for large-scale inference deployments.
AI Frameworks | Rust · llm-gateway · inference-routing
View Project →AWS CDK construct library providing multi-service, well-architected patterns for generative AI on AWS: Bedrock, SageMaker model deployment, RAG knowledge bases, OpenSearch vector stores, agents, and batch inference. Use it to define repeatable GenAI infrastructure in TypeScript, Python, Java, Go, or C#. Experimental, n
AI Frameworks | TypeScript · aws-cdk · infrastructure-as-code
View Project →OME is a Kubernetes operator for LLM serving: model-as-CRDs, GPU bin-packing scheduling, and automatic runtime selection across SGLang, vLLM, TensorRT-LLM and Triton. Adds prefill-decode disaggregation, multi-node inference, LoRA serving, autoscaling and benchmarking.
AI Frameworks | Go · kubernetes · llm-serving
View Project →Koishi chatbot plugin that adds multi-model LLM chat (OpenAI, Claude, Gemini, DeepSeek, Qwen, Ollama and more) via a LangChain-based adapter layer. Offers chat/browse/agent modes, YAML persona presets, MCP client tools, long-term memory, web search and text, voice or image output.
AI Frameworks | TypeScript · chatbot · koishi
View Project →ai4j is a JDK 8+ Java agentic SDK giving one API across OpenAI, Anthropic, DashScope, DeepSeek, Ollama and more, plus Tool Calling, MCP, A2A, RAG, Agent Runtime and a built-in Coding Agent CLI/TUI/ACP.
AI Frameworks | HTML · java · jdk8
View Project →Docs build pipeline and MDX source for docs.langchain.com, covering LangChain, LangGraph, LangSmith, and Deep Agents. It is a documentation monorepo, not an AI library or agent you install.
AI Frameworks | MDX · documentation · langchain
View Project →Curated awesome list of LLM pre-training resources: technical reports from Llama, Qwen, DeepSeek and others, training frameworks (Megatron-LM, DeepEP, DeepGEMM), open datasets and data-filtering methods. Useful as a reading and reference index for anyone pre-training or studying LLM training pipelines.
AI Frameworks | awesome-list · llm-pretraining
View Project →Official LangChain integration packages for Google AI: langchain-google-genai (Gemini API), langchain-google-vertexai (Vertex AI), and langchain-google-community. Provides Google chat models, embeddings, vector stores, and tools usable in any LangChain app.
AI Frameworks | Python · langchain · google
View Project →Monorepo of LangChain and LangGraph integrations for AWS: Bedrock/SageMaker LLMs, AWS vector stores and retrievers for RAG, Bedrock Agents and AgentCore tools, plus DynamoDB/Valkey checkpointers and memory stores. Successor to the AWS components in langchain-community.
AI Frameworks | Python · langchain · langgraph
View Project →Official WordPress AI plugin: a modular, opt-in framework that adds AI features (alt text, summarization, translation, image generation, comment moderation) to the Block Editor via the PHP AI Client and Abilities API, with connector plugins for OpenAI, Anthropic and Google.
AI Frameworks | PHP · wordpress · wordpress-plugin
View Project →B4.run is a TypeScript meta-framework that wraps LangGraph.js with file-system routes, generated route/state/tool types, workspace sandboxing, approval gates, durable threads, and fixture-backed tests to emit runnable Node servers and Dockerfiles.
AI Frameworks | TypeScript · agents · langgraph
View Project →PyTorch framework for continual learning of language models: implements DAS, CPT, DGA, EWC, HAT, DER++ and baselines for sequential domain-adaptive pretraining with end-task fine-tuning, forgetting-rate tools, and Hugging Face checkpoints.
AI Frameworks | Python · continual-learning · catastrophic-forgetting
View Project →ROS 2 packages that wrap llama.cpp and llava.cpp, exposing GGUF LLMs and VLMs as ROS 2 nodes with launch files, behavior-tree nodes, LangChain/RAG integration and LoRA/grammar support for local robotics inference.
AI Frameworks | C++ · ros2 · llama.cpp
View Project →Wavefront is an open-source middleware platform for building and operating enterprise AI agents, workflows, and RAG applications. It provides data integrations, MCP connectors, access controls, and observability.
AI Frameworks | Python · AI middleware · AI workflows
View Project →Agent Kernel is a Python platform layer for running, orchestrating and deploying production AI agents. It runs OpenAI Agents SDK, LangGraph, CrewAI and Google ADK side by side, adds guardrails, sessions, RAG, sandboxing, channels and MCP/A2A/AG-UI, and deploys to AWS, Azure, GCP or Kubernetes via Terraform and Helm.
AI Frameworks | Python · ai-agents · multi-agent
View Project →A Rust framework for building tool-using LLM agents, with streaming support across seven protocols, built-in tools, MCP and OpenAPI integrations, sub-agents, and session management. Includes a terminal coding-agent example.
AI Frameworks | Rust · agent loop · tool calling
View Project →