Awesome Data Engineering
Section: Stream Processing · An open source ETL framework to build fresh index for AI.
Entry
Appears in 6 awesome lists
ETL framework to build fresh context for AI agents, with incremental processing
Section: Stream Processing · An open source ETL framework to build fresh index for AI.
Section: Pipeline frameworks & libraries · ETL framework to build fresh index.
Section: Library
Section: CocoIndex (24 · 6.1K) - Data transformation framework for AI. Ultra performant, with.. Apache-2 · (👨💻 93 · 🔀 450 · 📦 73):
Section: MLOps · ETL framework to build fresh context for AI agents, with incremental processing
Section: LLM and Inference · Incremental engine for long horizon agents Star if you like it!
Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.
A fast and simple framework for building and running distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library. ray.io
Cloud-native orchestration platform for developing and maintaining data assets including ML models. Declarative programming model with integrated lineage and observability. Apache 2.0 licensed.
| Python | - Parallel computing with task scheduling in Python with a Pandas like API
| Python | - A scalable general purpose micro-framework for defining dataflows. You can use it to build dataframes, numpy matrices, python objects, ML models, etc. Embed Hamilton anywhere python runs, e.g. spark, airflow, jupyter, fastapi, python scripts, etc.
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG. Features 350+ connectors with always-in-sync data from SharePoint, Google Drive, S3, Kafka, PostgreSQL and more. BSL 1.1 license (becomes Apache 2.0 after 4 years).
An open-source embedding database for building AI applications with embeddings and semantic search.
is a library for efficient similarity search and clustering of dense vectors. It contains algorithms that search in sets of vectors of any size, up to ones that possibly do not fit in RAM. It also contains supporting code for evaluation and parameter tuning. Faiss is written in C++ with complete…