Vibe Coding Discover

Use Cases

Embeddings And Reranking For Rag Pipelines

Published projects tagged with this use case.

1 project

vllm-mlx

★ 1.6K

vLLM-style inference server for Apple Silicon built on MLX, exposing OpenAI /v1/* and Anthropic /v1/messages from one process. Adds continuous batching, paged/prefix KV cache, structured output, MCP tool calling, and multimodal text, vision, audio, embeddings and rerank support.

AI Frameworks | Python · llm-inference · mlx

View Project →