Vibe Coding Discover

RAG

Fast, Accurate, Lightweight Python library to make State of the Art Embedding

★ 3.2K262 forksPythonApache-2.0qdrant

FastEmbed is a lightweight Python library for generating dense, sparse, image, and late-interaction embeddings, plus reranking scores. It runs models with ONNX Runtime and supports CPU and GPU inference for retrieval pipelines.

Use Cases

Generate text embeddings for semantic searchCreate sparse embeddings for lexical retrievalEmbed images for multimodal searchGenerate late-interaction embeddings for document retrievalRerank search resultsBuild retrieval pipelines for QdrantRun embedding generation in lightweight serverless environments

Built With

Language
Python
Frameworks
ONNX Runtime · Hugging Face Hub · NumPy · Qdrant

Tags

embeddings · dense embeddings · sparse embeddings · reranking · vector search · retrieval · multimodal · ONNX · CPU inference · GPU inference