ActionChain
A workflow system for simple linear success/failure workflows.
A curated list of awesome pipeline toolkits inspired by Awesome Sysadmin
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
A workflow system for simple linear success/failure workflows.
Small package to describe workflows that are not completely known at definition time.
workflow manager with a strong focus on provenance, performance and extensibility.
Python-based workflow system created by AirBnb.
Component-based workflow framework for scientific data analysis.
High-level language for biology.
Container-native workflow engine for orchestrating parallel data processing, ML, or CI jobs on Kubernetes.
An open source Python experiment and workflow manager used to manage complex workflows on Cloud and HPC platforms.
Workflow and resource management system with CWL support.
Python-based high throughput task and workflow engine.
Scripting language for data pipelines.
Unified programming model for batch and streaming data-parallel processing pipelines.
GNU-Make-like utility for managing builds and complex workflows.
Explicit framework with web monitoring and resource estimation.
Haskell DSL built on shake with strong typing and EDAM support.
Library to build and execute typed scientific workflows.
Tool for running and managing bioinformatics pipelines.
Python Meta-programming Library for Job Flow Control.
Python based lightweight graph (i.e. can do loops and conditional branching, and not just DAGs) orchestrator.
Command-line tool which uses common cluster managers to run bioinformatics pipelines.
Automated reproducibility, and hassle-free submission of computational jobs to clusters.
Application framework for portable computational pipelines.
Programming model for distributed infrastructures.
Light-weight workflow management application.
A Python pipeline abstraction inspired by Apache Storm topologies.
Python library for massively parallel workflows.
Unified interface for constructing and managing workflows on different workflow engines, such as Argo Workflows, Tekton Pipelines, and Apache Airflow.
Workflow orchestration toolkit for high-performance and quantum computing research and development.
Workflow Management System geared towards scientific workflows from the Broad Institute.
Advanced functional workflow language and framework, implemented in Erlang.
A workflow engine for cycling systems, originally developed for operational environmental forecasting.
Simple DAG-based job scheduler in Python.
A scala based DSL and framework for writing and executing bioinformatics pipelines as Directed Acyclic Graphs.
Python-based API for defining DAGs that interfaces with popular workflow managers for building data applications.
an open-source relational framework for scientific data pipelines.
Framework for writing analytics workflows entirely in SQL. The T part of ETL, focuses on analytics engineering.
Workflow runner that uses Dataflow to run a series of tasks in Docker.
Python library for creating pipelines and workflows easily.
Robust DSL akin to Make, implemented in Clojure.
Reproducibility and high-performance computing with an easy R-focused interface. Unrelated to Factual's Drake. Succeeded by Targets.
An engine for managing the execution of container-based workflows.
Workflow manager.
System for creating and running pipelines on a distributed compute resource.
A fast, lightweight workflow engine for serverless/FaaS functions.
Language agnostic framework for building flexible data science pipelines (Python/Shell/Gnuplot).
Robust and efficient workflows using a simple language agnostic approach (R package).
Python libraries and tools for running applications on diverse Grids and clusters.
A workflow management language extension for GNU Guix.
Make-like utility for submitting workflows via qsub.
A python micro-framework for describing dataflows; runs anywhere python runs.
Hera is an Argo Python SDK. Hera aims to make construction and submission of various Argo Project resources easy and accessible to everyone! Hera abstracts away low-level setup details while still maintaining a consistent vocabulary with Argo.
Platform for defining and executing workflow pipelines in large-scale distributed environments.
HPC-focused task scheduler that automatically assigns tasks to Slurm/PBS allocations and submits them for the user.
Set of tools to provide lightweight pipelining in Python.
A task Based parallelization framework for Python.
Open source data orchestration and scheduling platform with declarative syntax.
Embedded DSL in the OCAML language alongside a client-server management application.
Python framework for building efficient data pipelines.
Workflow assembler for cancer genome analytics and informatics.
Framework for building and deploying portable, scalable machine learning workflows using Docker containers and Argo Workflows.
Tool for running bioinformatics workflows locally or in the cloud.
Job proxying tool for biomolecular simulations.
YAML based HPC workflow execution tool.
Workflow engine for executing large complex workflows on clusters.
Run R scripts if needed, based on last modified time.
An R package which provides a set of simple tools for transforming an existing workflow into a self-documenting pipeline with very minimal upfront costs.
A lightweight, opinionated ETL framework, halfway between plain scripts and Apache Airflow.
Scala library for defining data pipelines.
A language and framework for developing and executing complex computational pipelines.
Microservice based workflow engine.
Open-sourced framework from Netflix, for DAG generation for data scientists. Python and R API's.
Python based workflow engine by the Open Stack project.
Lightweight workflows in bioinformatics.
Flow-based computational toolkit for reproducible and scalable bioinformatics pipelines.
Embeddable JVM-based workflow engine with high availability, fault tolerance, and support for multiple databases. Additional libraries are provided for visualization and REST API.
Workflows and interfaces for neuroimaging packages.
Accelerated framework for manipulating and interpreting high-throughput sequencing data.
Distributed and reproducible data pipelining and data management, built on the container ecosystem.
Productive parallel programming, for creating parallel programs composed of Python functions and external components.
Build pandas dataframe processing pipelines interactively in notebooks and save for command line resusage. Generally applicable but currently focuses on chemistry (RDKit).
Lightweight function pipeline (DAG) creation in pure Python for scientific workflows.
Ruby based launcher for complex biological pipelines.
Python based workflow engine by Pinterest.
YAML based container-native workflow engine supporting Docker, Singularity, Vagrant VMs with Docker daemon in VM, and local host.
Haskell workflow tool to express and compose tasks (optionally cached) whose datasources and sinks are known ahead of time and rebindable, and which can expose arbitrary sets of parameters to the outside world.
Python based workflow engine powering Prefect.
Lightweight, DAG-based Python dataflow engine for reproducible and scalable scientific pipelines.
Lightweight parallel task engine.
Simple push-based python workflow framework using asyncio, supporting recursive networks.
A python lightweight pipeline framework.
Automation task-runner for sequential steps defined in a pipeline yaml, with AWS and Slack plug-ins.
A workflow management system that facilitates reproducible data analyses.
Parallel workflow extension for Rake.
Lightweight high-throughput queuing system for workflows with many small tasks to perform.
Simple tokenised template system for SGE.
Python-based workflow toolkit based on the Common Workflow Language and Docker.
Framework for large distributed task-based pipelines, written in Rust with Python API.
Yet another redundant workflow engine.
Language and runtime for distributed, incremental data processing in the cloud.
Make-like declarative workflows in R.
Local-first IDE for authoring and running SQL and Python data pipelines stored in Git.
Wrapper for the creation of Makefiles, enabling massive parallelization.
Pipeline system for bioinformatics workflows.
Computation Pipeline library for Python.
Pipeline tool for R, inspired by Luigi.
Language and runner for real-time scheduling and logistics.
Self-documenting build automation tool.
Helper library for writing flexible scientific workflows in Luigi.
Library for writing Scientific Workflows in Go.
Lightweight, but scalable framework for file-driven workflows to be run locally and on HPC systems.
Scalable Concurrent Operations in Python.
Python library for lazy evaluation of pipelined transformations on indexable containers.
A framework for rapid development of robust data pipelines following a simple design pattern.
Tool for running and managing bioinformatics pipelines.
Based on the Workflow Patterns initiative and implemented in Python.
Directed Acyclic Graph task dependency scheduler that simplify distributed pipelines.
lightweight, open-source, Python 3 library for fast and reproducible experimentation. (This repository has been archived by the owner on Jun 22, 2022.)
File processing pipelines as a Python library.
Container native workflow management system focused on hybrid workflows.
A self-service IoT toolbox to enable non-technical users to connect, analyze and explore IoT data streams.
Jobsystem on AWS ECS or AWS Batch managing dependencies and scheduling.
Java-based distributed pipeline from Netflix.
Fast easy parallel scripting - on multicores, clusters, clouds and supercomputers.
R package to organize reproducible scientific workflows.
Dynamic, function-oriented Make-like reproducible pipelines at scale in R.
A library to help manage complicated computational software pipelines consisting of long running individual tasks.
Tool that helps you run genomic pipelines on Amazon cloud.
Distributed pipeline workflow manager (mostly for genomics).
Workflow compute layer that lets any job start instantly, scale across clouds, and avoid dependency drift by design.
Extensible parallel framework, written in Python using OpenMPI libraries.
A C++ parallel pipeline library for stream processing.
Framework for streaming data applications and algorithms that react to real-time events.
An open source implementation of a Workflow Execution Service (WES). It tries to be GA4GH-compliant.
Easy Collaborative Reproducible Computing.
Workflow engine for orchestrating jobs, data and events across your applications and third party services.
Extensible open-source MLOps framework to create reproducible pipelines for data scientists.
Computational science made reproducible and publishable.
Polyglot workflows without leaving the comfort of your technology stack.
A community and framework centered around metagenomics, designed to facilitate reproducible exploration and visualization of data.
Framework for executing and managing computational workflows on distributed computing resources.
Event-driven automation for sequencing centers. Initiates workflows based on events.
A container based workflow platform.
Framework for running scientific workflows on public and academic clouds.
Open source platform for data analysis.
Cluster Load Balancer for Bioinformatics e-Resources.
Workflow manager designed for simplicity, extensibility and collaboration.
User friendly and open source visual workflow management platform.
Centralized workflow server for dynamic workflows of high-throughput computations.
Open source visual Python scripting for test, measurement, and robotics control.
Open-source local data platform with a visual pipeline builder and a Polars-compatible Python API.
FlowHub is a new workflow cloud platform.
Container-native, type-safe workflow and pipelines platform for large scale processing and ML.
Powerful workflow system which can be used on the command line or with the GUI.
In-browser tool for data processing workflows with high-performance server support, featuring code history and workflow orchestration.
Kepler scientific workflow application from University of California.
General-purpose platform with many specialized domain extensions.
Toolkit for making deployments of machine learning workflows on Kubernetes simple, portable and scalable.
Data & model pipeline deployment for humans - integrated, scalable, extensible.
Workflow Management System for exploration of models and parameter optimization.
Data-analytics platform with declarative workflows of distributed operations.
Workflow Management System.
Distributed workflow engine designed to be dead simple.
Platform for reusable research data analyses developed by CERN.
Supporting User for SHell script Integration.
Domain independent workflow system.
Highly scalable developer oriented Workflow as Code engine.
Workflow compute layer that lets any job start instantly, scale across clouds, and avoid dependency drift by design.
Scientific workflow and provenance management system.
Workflow management system for the automated and distributed analysis of large-scale experimental data.
Developer platform and workflow engine to turn scripts into internal tools.
Semantic workflow system utilizing Pegasus as execution system.
Online research environment for grid, HPC and cloud computing.
a specification for describing analysis workflows and tools that are portable and scalable across a variety of software and hardware environments, from workstations to cluster, cloud, and high performance computing (HPC) environments. [ web ]
Workflows across research domains and languages
Containers and workflows, in particular life sciences
Curated Nextflow workflows for life sciences
(IWC) - Curated Galaxy workflows
git and git-annex based data version control system with lightweight provenance capture/re-execution support.
Data version control system for ML project with lightweight pipeline support.
Provides Git-like capability & version control for Iceberg Tables, Delta Lake Tables & SQL Views.
Real-time data quality screening API that returns PASS / WARN / BLOCK verdicts at the ingest boundary before data enters pipelines or warehouses in milliseconds.Python SDK available.
Notebook-style development environment.
Turn a GitHub repo into a collection of interactive notebooks powered by Jupyter and Kubernetes
A rich architecture for interactive computing.
GNU Emacs major mode for computational notebooks, literate programming, and much more.
Interactive data workflows built on Python.
A better notebook for Scala (and more). Built by Netflix.
Consolidate your notebooks and scripts in a reproducible pipeline using a pipeline.yaml file
R Markdown notebook literate programming environment.
Web-native computational notebook for programmers supporting multiple languages, APIs and webooks.
Readable, interactive, cross-platform and cross-language data science workflow system.
Driver for ADBC, Apache Arrow's database connectivity API, that runs on top of any ODBC driver, so databases shipping ODBC but no Arrow driver (Db2, Informix, Vertica, Firebird, Ingres and more) return columnar Arrow batches and take bulk loads. One C library with Python, Rust, Go, Java and C#…
Data pipeline framework supporting SQL and Python in the same DAG. Built-in data quality assertions, cross-database lineage, and incremental processing. Targets data warehouses (BigQuery, Snowflake, Postgres, etc.).
Distributed, scalable, durable, and highly available orchestration engine developed by Uber.
Dataform is a framework for managing SQL based operations in your data warehouse.
Self-hosted ELT platform combining dlt extract-and-load, dbt-core transforms and scheduling in one UI.
Hevo is a Fully Automated, No-code Data Pipeline Platform that supports 150+ ready-to-use integrations across Databases, SaaS Applications, Cloud Storage, SDKs, and Streaming Services.
Declarative ETL engine where pipelines are YAML files validated before they run.
A data processing & ETL framework for Ruby.
Linked Data publishing and consumption ETL tool.
Polyglot data loader based on dlt for ETL, warehousing and more. Copy data between any source and any destination. Supports 140+ source/destination adapters. Fast transformations based on Polars expressions.
Performant open-source Python ETL framework with Rust runtime, supporting 300+ data sources.
A plataform that delivers poweful ETL capabilities, using a groundbreaking, metadata-driven approach.
Substation is a cloud native data pipeline and transformation toolkit written in Go.
Managed cloud object storage transfers for ingestion workflows.
A tool for the automated exploration of possible computational workflows based on semantic annotations.
Powerful and scalable directed graphs of data routing, transformation, and system mediation logic.
Supporting infrastructure to run scientific experiments without a scientific workflow management system, and still get things like provenance.
Simplifies the process of creating reproducible experiments from command-line executions.
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
ai-boost/awesome-prompts
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
matiassingers/awesome-readme
A curated list of awesome READMEs