Skip to content
72

Awesome ETL

A curated list of awesome ETL frameworks, libraries, and software.

3.6k stars376 forks65 entriesLast push May 1, 2026 (5 months ago)License CC0-1.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Workflow Management/Engines

Airflow

"Use airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command line utilities make performing complex surgeries on DAGs a snap. The rich user interface makes it…

In 13 listsDetails

Argo

"an open source container-native workflow engine for orchestrating parallel jobs on Kubernetes."

In 2 lists

Dagster

"Dagster is a data orchestrator for machine learning, analytics, and ETL. It lets you define pipelines in terms of the data flow between reusable, logical components, then test locally and run anywhere. With a unified view of pipelines and the assets they produce, Dagster can schedule and…

In 2 lists

Luigi

"a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in."

In 15 listsDetails

prefect

"a workflow orchestration framework for building resilient data pipelines in Python."

In 11 listsDetails

Temporal

"a scalable and reliable runtime for durable function executions called Temporal Workflow Executions."

In 3 lists

Toil

"an open-source pure-Python workflow engine that lets people write better pipelines."

Job Scheduling

Jenkins

"the leading open-source automation server. Built with Java, it provides over 1000 plugins to support automating virtually anything, so that humans can actually spend their time doing things machines cannot."

In 5 listsDetails

Java

Apache Camel

"an open source integration framework that empowers you to quickly and easily integrate various systems consuming or producing data."

In 4 listsDetails

Spring Batch

"A lightweight, comprehensive batch framework designed to enable the development of robust batch applications that are vital for the daily operations of enterprise systems."

Python >Libraries

BeautifulSoup

"a Python library for pulling data out of HTML and XML files."

Celery

"an asynchronous task queue/job queue based on distributed message passing. It is focused on real-time operation, but supports scheduling as well."

Dask

"a flexible parallel computing library for analytics."

In 11 listsDetails

dataset

A wrapper around SQLAlchemy that simplifies database operations (including upserting).

dbt-core

"enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications."

In 7 listsDetails

dlt

"an open-source Python library that loads data from various, often messy data sources into well-structured datasets."

DuckDB

"an analytical in-process SQL database management system."

In 8 listsDetails

Great Expectations

"a Python library for validating, documenting, and profiling your data to maintain quality and improve communication between teams about data and data pipelines."

hamilton

"helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does."

In 10 listsDetails

ijson

"Iterative JSON parser with Pythonic interfaces."

ingestr

"a CLI tool to copy data between any databases with a single command seamlessly."

In 6 listsDetails

Joblib

"a set of tools to provide lightweight pipelining in Python."

lxml

"the most feature-rich and easy-to-use library for processing XML and HTML in the Python language."

In 3 lists

Meltano

"the declarative code-first data integration engine."

In 3 lists

Pandas

"Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more."

In 2 lists

parse

"Parse strings using a specification based on the Python format() syntax."

PETL

"a general purpose Python package for extracting, transforming and loading tables of data."

In 3 lists

polars

"Extremely fast Query Engine for DataFrames, written in Rust."

In 10 listsDetails

PyQuery

"A jquery-like library for python."

Scrapy

"a fast high-level web crawling & scraping framework for Python."

In 5 listsDetails

SQLAlchemy

"the Python SQL toolkit and Object Relational Mapper that gives application developers the full power and flexibility of SQL."

tenacity

"a general-purpose retrying library, written in Python, to simplify the task of adding retry behavior to just about anything."

In 3 lists

Toolz

"A functional standard library for Python."

xmltodict

"Python module that makes working with XML feel like you are working with JSON."

In 5 listsDetails

Ruby

Embulk

"a parallel bulk data loader that helps data transfer between various storages, databases, NoSQL and cloud services."

In 4 listsDetails

Kiba

"lets you define and run high-quality ETL jobs using Ruby."

In 2 lists

nokogiri

"Nokogiri makes it easy and painless to work with XML and HTML from Ruby."

In 5 listsDetails

Sequel

"a simple, flexible, and powerful SQL database access toolkit for Ruby."

In 2 lists

Go

CloudQuery

"a cloud asset inventory built for platform teams. Sync your cloud infrastructure metadata into your data warehouse, powering insights and automation."

In 6 listsDetails

Pachyderm

"provides parallelized processing of multi-stage, language-agnostic pipelines with data versioning and data lineage tracking."

In 5 listsDetails

Redpanda Connect

"a declarative data streaming and integration tool with 300+ pre-built connectors, configured via YAML."

Cloud Services

Airbyte

"Airbyte is an open-source data integration engine that helps you consolidate your data in your data warehouses, lakes and databases."

In 2 lists

Alteryx

"combines data preparation, data blending, and analytics — predictive, statistical, and spatial — in a visual workflow designer."

AWS Batch

"enables developers, scientists, and engineers to easily and efficiently run hundreds of thousands of batch computing jobs on AWS."

In 2 lists

AWS Glue

"a serverless data integration service that makes it easy for analytics users to discover, prepare, move, and integrate data from multiple sources."

In 3 lists

Cloud Data Fusion

"Fully managed, cloud-native data integration platform."

Fivetran

"automates data movement from disparate sources into your destination."

In 3 lists

Google Dataflow

"Google Cloud Dataflow provides a simple, powerful model for building both batch and streaming parallel data processing pipelines."

Hevo

"a no-code data movement platform that is usable by your most technical as well as your non-technical and business users."

In 4 lists

Microsoft Azure Data Factory

"A fully managed, serverless data integration service that helps you visually integrate data sources with more than 90 built-in, maintenance-free connectors."

Stitch

"Stitch is a cloud-first, open source platform for rapidly moving data. A simple, powerful ETL service, Stitch connects to all your data sources – from databases like MySQL and MongoDB, to SaaS applications like Salesforce and Zendesk – and replicates that data to a destination of your choosing."

In 3 lists

Big Data (Hadoop Stack)

Apache Beam

"a unified programming model for Batch and Streaming data processing."

In 6 listsDetails

Apache Flink

"a framework and distributed processing engine for stateful computations over unbounded and bounded data streams."

In 6 listsDetails

Debezium

"Change data capture for a variety of databases."

In 2 lists

Kafka Connect

"a tool for scalably and reliably streaming data between Apache Kafka and other systems. It makes it simple to quickly define connectors that move large collections of data into and out of Kafka."

In 4 lists

Spark

"a fast and general-purpose cluster computing system. It provides high-level APIs in Scala, Java, and Python that make parallel jobs easy to write, and an optimized engine that supports general computation graphs. It also supports a rich set of higher-level tools including Shark (Hive on Spark),…

In 7 listsDetails

ETL Tools (GUI)

Apache NiFi

"a rich, web-based interface for designing, controlling, and monitoring a dataflow."

In 5 listsDetails

CDAP

"Use Cask Data Application Platform to visually build and manage data applications in hybrid and multi-cloud environments."

Informatica PowerCenter

An ETL tool for extracting data from source systems, transforming it, and loading it into target systems using a visual mapping and workflow designer.

In 2 lists

Microsoft SSIS

"a component of the Microsoft SQL Server database software that can be used to perform a broad range of data migration tasks."

N8n

"Free and open fair-code licensed node based Workflow Automation Tool. Easily automate tasks across different services."

In 9 listsDetails

Pentaho Data Integration (PDI)

"a graphical ETL tool for designing data integration workflows using a drag-and-drop interface, also known as Kettle."

Further Reading

Fundamentals of Data Engineering

Joe Reis & Matt Housley's tool-agnostic overview of the data engineering lifecycle, including the ETL-to-ELT shift (2022).

The Rise of Data Contracts

Chad Sanderson on formalizing schema and quality guarantees between data producers and consumers.

ELT 101: The Why and What of ELT

Why cheap cloud warehouse compute flipped the ETL paradigm to ELT.

See category
87

Awesome Data Engineering

igorbarinov/awesome-data-engineering

A curated list of data engineering tools for software developers

Fresh★ 9.1k313 entriesPushed 22 days ago
84

Awesome Big Data

oxnr/awesome-bigdata

A curated list of awesome big data frameworks, ressources and other awesomeness.

Active★ 15k645 entriesPushed 2 months ago
84

Awesome Public Datasets

awesomedata/awesome-public-datasets

A topic-centric list of HQ open datasets.

Fresh★ 79k3 entriesPushed today
83

Awesome Ada

ohenley/awesome-ada

A curated list of awesome resources related to the Ada and SPARK programming language

Fresh★ 869418 entriesPushed 7 days ago
80

Awesome Network Analysis

briatte/awesome-network-analysis

A curated list of awesome network analysis resources.

Fresh★ 4.1k717 entriesPushed 1 month ago
78

Awesome Public Real-Time Datasets and Sources

bytewax/awesome-public-real-time-datasets

A list of publicly available datasets with real-time data maintained by the team at bytewax.io

Active★ 2.9k91 entriesPushed 2 months ago