Skip to content

Entry

Pachyderm

Appears in 5 awesome lists

Open source distributed processing framework build on Kubernetes focused mainly on dynamic building of production machine learning pipelines - (Video).

Open github.compachyderm/pachyderm

Found in these lists

Awesome Cloud Native

Section: Data Processing & Analytics · Reproducible Data Science at Scale!

FreshScore 87

Awesome ETL

Section: Go · "provides parallelized processing of multi-stage, language-agnostic pipelines with data versioning and data lineage tracking."

ActiveScore 72

Awesome LLMOps

Section: Data Management · Pachyderm is a version control system for data.

ActiveScore 75

Awesome Production Machine Learning

Section: Data Pipeline · Open source distributed processing framework build on Kubernetes focused mainly on dynamic building of production machine learning pipelines - (Video).

FreshScore 92

awesome-go

Section: Data Science and Analytics · Data-Centric Pipelines and Data Versioning

FreshScore 81

Dolt

Dolt is Git for Data! Dolt is a SQL database that you can fork, clone, branch, merge, push and pull just like a git repository.

In 10 listsDetails

DVC

Version control for large files. kedro - Build data pipelines. feast - Feature store. Video. pgvector - Vector similarity search for Postgres. pinecone - Database for vector search applications. truss - Serve ML models. milvus - Vector database for similarity search. mlem - Version and deploy your…

In 8 listsDetails

CloudQuery

Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.

In 6 listsDetails

Delta Lake

Open-source storage framework enabling Lakehouse architecture with ACID transactions, scalable metadata handling, and unified batch/streaming processing. Apache 2.0 licensed.

In 6 listsDetails

SQLFlow

Brings machine learning capabilities to SQL, enabling model training and prediction using SQL syntax.

In 6 listsDetails

Quilt

Versioning, reproducibility and deployment of data and models.

In 3 lists

wallaroo

Ultrafast and elastic data processing.

fast-data-dev

Kafka Docker for development. Kafka, Zookeeper, Schema Registry, Kafka-Connect, Landoop Tools, 20+ connectors.