self-hosted-ai-stack
★ 156Docker Compose bundle that deploys a full local AI stack: Ollama for LLMs, LiteLLM gateway, AnythingLLM chat UI, embeddings/RAG, Whisper STT, Kokoro TTS, Docling parsing, and an MCP Gateway. Includes lightweight stack variants, optional HTTPS and CUDA GPU acceleration.
AI Frameworks | Shell · self-hosted · docker-compose
View Project →swift-ai-sdk
★ 154Swift port of the Vercel AI SDK offering one provider-agnostic API for streaming text, structured outputs, tool calling, MCP tools, and middleware across 38 providers via SwiftPM. Built for iOS/macOS apps needing OpenAI, Anthropic, Google and others from Swift.
AI Frameworks | Swift · sdk · llm
View Project →neurolink
★ 144TypeScript AI SDK unifying 30+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Ollama) behind one streaming API. MCP-native with built-in RAG, memory, voice TTS/STT, agents, and provider failover. Extracted from Juspay production systems.
AI Frameworks | TypeScript · llm · multi-provider
View Project →llm
★ 141llm.rb is a zero-dependency Ruby runtime for building agentic LLM apps: a single API across 14+ providers, managed tool loops, tools, skills, MCP/A2A clients, streaming callbacks, concurrency strategies and ActiveRecord/Sequel persistence, plus an interactive agent console.
AI Frameworks | Ruby · llm-runtime · agents
View Project →model-compose
★ 113model-compose is a declarative YAML orchestrator (docker-compose for AI) that deploys chat APIs, ReAct agents, RAG pipelines, and MCP servers from one file. It bridges local and cloud models and runs on Docker, native, or distributed Redis-queued runtimes.
AI Frameworks | Python · yaml · declarative
View Project →ai-microcore
★ 108MicroCore is a minimalist Python library of LLM and vector-DB adapters that makes providers switchable via config while keeping app code unchanged. It includes prompt templating, streaming, embeddings search (Chroma/Qdrant) and LLM-agnostic MCP tool integration.
AI Frameworks | Python · llm · mcp
View Project →quarkus-workshop-langchain4j
★ 107A step-by-step Java workshop for building AI applications with Quarkus and LangChain4j, covering single AI services and agentic orchestration. Each lesson has a runnable project state.
AI Frameworks | Java · workshop · Quarkus
View Project →HOMER
★ 45Official ICLR 2024 implementation of HOMER, a training-free hierarchical KV-cache merging method that extends pre-trained LLM context limits (e.g. Llama-2) with lower memory. Ships patched LlamaForCausalLM, plus passkey-retrieval and PG19 perplexity scripts.
AI Frameworks | Python · long-context · kv-cache
View Project →