Skip to content

Entry

Apache Flume

Appears in 4 awesome lists

Apache Flume is a distributed, reliable, and available system for efficiently collecting, aggregating and moving large amounts of log data from many different sources to a centralized data store. License: Apache 2.

Open flume.apache.org

Found in these lists

Awesome Big Data

Section: Data Ingestion · service to manage large amount of log data.

ActiveScore 84

Awesome Cloud Native

Section: Logging · Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data.

FreshScore 87

Awesome Hadoop

Section: Data Ingestion and Integration · Apache Flume

StaleScore 48

Useful Java links

Section: 7. Big data · Apache Flume is a distributed, reliable, and available system for efficiently collecting, aggregating and moving large amounts of log data from many different sources to a centralized data store. License: Apache 2.

ActiveScore 74

Telegraf

Input Plugin for Telegraf to collect and forward Suricata stats logs (included out of the box in recent Telegraf releases).

In 8 listsDetails

Elasticsearch

Distributed search and analytics engine with native k-NN vector search, hybrid search, and dense vector indexing. Industry-standard for full-text search now with powerful semantic search capabilities. AGPL-3.0/Elastic-2.0 dual licensed.

In 7 listsDetails

Logstash

Ship logs from any source, parse them, get the right timestamp, index them, and search them. Logstash is a tool for managing events and logs. You can use it to collect logs, parse them, and store them for later use (like, for searching). Speaking of searching, Logstash comes with a web interface…

In 8 listsDetails

Apache Pulsar

a distributed pub-sub messaging platform with a very flexible messaging model and an intuitive client API.

In 6 listsDetails

vector

A high-performance observability data pipeline for collecting, transforming, and routing logs and metrics. Real-time data processing with 50+ sources and sinks including Kafka, S3, and Elasticsearch. Ideal for AI/ML log processing and data ingestion. MPL 2.0 licensed.

In 5 listsDetails

ingestr

CLI tool to copy data from any source to any destination with a single command. Load data into your warehouse before dbt transforms it. Supports 50+ sources including Postgres, MongoDB, Salesforce, Shopify.

In 6 listsDetails

Bruin

Run and schedule dbt-style SQL transformations without Airflow. Adds data ingestion (50+ sources) and built-in data quality to the transformation layer. Open-source CLI or managed Bruin Cloud for teams who want dbt Cloud-like experience with ingestion included.

In 6 listsDetails

Kafka

🛠 - 🐙 - Distributed, fault tolerant, high throughput pub-sub messaging system.

In 6 listsDetails