Awesome Data Analysis
Section: Useful Python Tools for Data Analysis · A lightweight pipelining library for Python, particularly useful for saving and loading large NumPy arrays.
Entry
Appears in 5 awesome lists
A lightweight pipelining library for Python, particularly useful for saving and loading large NumPy arrays.
Section: Useful Python Tools for Data Analysis · A lightweight pipelining library for Python, particularly useful for saving and loading large NumPy arrays.
Section: Distributed Computing · Parallel computing and disk-based caching for Python functions.
Section: joblib (34 · 4.4K · ) - Computing with Python functions. BSD-3 · (👨💻 160 · 🔀 470 · 📦 750K):
Section: Concurrency and Parallelism · Computing with Python functions.
Section: Computation · | Python | - running Python functions as pipeline jobs
A fast and simple framework for building and running distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library. ray.io
| Python | - Parallel computing with task scheduling in Python with a Pandas like API
| Rust, Python | - Polars is a blazingly fast DataFrames library implemented in Rust using Apache Arrow Columnar Format as memory model.
(label: good first issue) Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
Vaex is a high performance Python library for lazy Out-of-Core DataFrames (similar to Pandas), to visualize and explore big tabular datasets. Vaex uses memory mapping, zero memory copy policy and lazy computations for best performance (no memory wasted).
is an open source, NumPy-aware optimizing compiler for Python sponsored by Anaconda, Inc. It uses the LLVM compiler project to generate machine code from Python syntax. Numba can compile a large subset of numerically-focused Python, including many NumPy functions. Additionally, Numba has support…
Unified analytics engine for large-scale data processing. In-memory cluster computing with high-level APIs in Python, Scala, Java, and R. Powers MLlib for distributed machine learning and Structured Streaming for real-time data. Apache 2.0 licensed.