Awesome Big Data
Section: Distributed Programming · an unified model and set of language-specific SDKs for defining and executing data processing workflows.
Entry
Appears in 6 awesome lists
Unified data processing engine supporting both batch and streaming applications. Apache Spark is one of the supported execution environments.
Section: Distributed Programming · an unified model and set of language-specific SDKs for defining and executing data processing workflows.
Section: General · by Apache
Section: Stream Processing · A unified programming model that implements both batch and streaming data processing jobs that run on many execution engines.
Section: Big Data (Hadoop Stack) · "a unified programming model for Batch and Streaming data processing."
Section: Pipeline frameworks & libraries · Unified programming model for batch and streaming data-parallel processing pipelines.
Section: Interfaces · Unified data processing engine supporting both batch and streaming applications. Apache Spark is one of the supported execution environments.
Taskade AI is an AI-powered productivity suite offering tools like task and project management, notes, docs, mind maps, and AI chat to enhance team productivity and automate over 700 tasks website | github | twitter | youtube
write, plan, collaborate, and get organized. Notion is all you need — in one tool.
Python module for building complex pipelines of batch jobs. Handles dependency resolution, workflow management, visualization, and Hadoop integration. Built at Spotify and battle-tested in production. Apache 2.0 licensed.
A fast and simple framework for building and running distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library. ray.io
🆓 Connect and query your data sources, build dashboards to visualize data and share them with your company. Owned by Databricks but the hosted SaaS shut down in 2021, so the project is now community-maintained under the Databricks org with no paid Redash product.
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG. Features 350+ connectors with always-in-sync data from SharePoint, Google Drive, S3, Kafka, PostgreSQL and more. BSL 1.1 license (becomes Apache 2.0 after 4 years).
is a professional DataBase IDE developed by Jet Brains that provides context-sensitive code completion, helping you to write SQL code faster. Completion is aware of the tables structure, foreign keys, and even database objects created in code you're editing.
is a powerful, open source object-relational database system with over 30 years of active development that has earned it a strong reputation for reliability, feature robustness, and performance.