Awesome Data Engineering
Section: Data Lake Management · An open source platform that delivers resilience and manageability to object-storage based data lakes.
Entry
Appears in 10 awesome lists
Data version control for your data lake that transforms object storage into Git-like repositories. Enables atomic, versioned data lake operations with branching, committing, and merging for data pipelines. Apache 2.0 licensed.
Section: Data Lake Management · An open source platform that delivers resilience and manageability to object-storage based data lakes.
Section: Data Storage · Git-like capabilities for your object storage.
Section: Data Management · Repeatable, atomic and versioned data lake on top of object storage.
Section: 1. Core Frameworks & Libraries · Data version control for your data lake that transforms object storage into Git-like repositories. Enables atomic, versioned data lake operations with branching, committing, and merging for data pipelines. Apache 2.0 licensed.
Section: General · ] - Repeatable, atomic and versioned data lake on top of object storage.
Section: ETL & Data orchestration · Repeatable, atomic and versioned data lake on top of object storage.
Section: Model, Data and Experiment Management · Repeatable, atomic and versioned data lake on top of object storage.
Section: S3 compatible file servers · lakeFS is an open source tool that transforms your object storage into a Git-like repository. It enables you to manage your data lake the way you manage your code.
Section: 文件/存储 · 类 Git 文件对象存储
Section: Data Science and Analytics · lakeFS - Data version control for your data lake | Git for data
Milvus is a cloud-native, open-source vector database built to manage embedding vectors generated by machine learning models and neural networks.
Dolt is Git for Data! Dolt is a SQL database that you can fork, clone, branch, merge, push and pull just like a git repository.
A high-performance Cloud-Native file system driven by object storage for large-scale data storage.
Vector Search Engine and Database for the next generation of AI applications. Also available in the cloud
The Pinecone vector database makes it easy to build high-performance vector search applications. Developer-friendly, fully managed, and easily scalable without infrastructure hassles.
Open-source storage framework enabling Lakehouse architecture with ACID transactions, scalable metadata handling, and unified batch/streaming processing. Apache 2.0 licensed.
LF AI & Data Foundation Graduated project for metadata collection, aggregation, and visualization. Maintains provenance of how datasets are consumed and produced with global visibility into job runtime and dataset lifecycle management. Integrates with OpenLineage. Apache 2.0 licensed.