Open-source inference server and production cluster for all the models your agent needs.
Self-hosted inference server and production cluster for agent workloads. Serves embedding, retrieval, reranking, generation, OCR, extraction, and multimodal models through an OpenAI-compatible API, with on-demand model loading and integrations for popular AI frameworks and vector stores.
Use Cases
Serve agent models from a self-hosted clusterGenerate embeddings for semantic searchRerank retrieved documentsConvert PDFs and office documents to MarkdownExtract entities and structured dataRun chat and text completion APIsClassify and moderate contentTranscribe audio and analyze images
Built With
- Language
- Python
- Frameworks
- LangChain · LlamaIndex · Haystack · DSPy · CrewAI · Chroma · Qdrant · Weaviate · LanceDB · SGLang
Tags
inference server · model serving · agent infrastructure · embeddings · reranking · semantic search · RAG · OpenAI-compatible API · self-hosted · Kubernetes · multimodal