datachain
View on GitHubThe Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Python library for building typed, versioned datasets from unstructured files in cloud storage. It supports incremental processing, checkpoint recovery, metadata and vector queries, plus an optional knowledge base and agent harness for coding agents.
Use Cases
Index files in S3, GCS, Azure, or local storage into typed datasetsBuild resumable pipelines for metadata extraction and embedding generationSearch and filter embeddings alongside structured metadataCreate a markdown knowledge base of datasets for humans and coding agentsGive coding agents access to versioned data through skills and MCP
Built With
- Language
- Python
- Frameworks
- Pydantic · Pandas · PyArrow · SQLAlchemy · LiteLLM · DVC · PyTorch · Transformers · OpenCLIP
Tags
unstructured-data · agent-context · versioned-datasets · data-pipelines · knowledge-base · vector-search · embeddings · cloud-storage · multimodal · incremental-processing · data-lineage · agent-harness