Articles about late data on waitingforcode.com

January 9, 2019 • Apache Spark Structured Streaming

Apache Spark 2.4.0 features - watermark configuration

The series about Apache Spark 2.4.0 features continues. After last week's discovery of bucket pruning, it's time to switch to Structured Streaming module and see its major evolution.

Continue Reading →

March 4, 2018 • Apache Spark Structured Streaming

Apache Spark Structured Streaming and watermarks

The idea of watermark was firstly presented in the occasion of discovering the Apache Beam project. However it's also implemented in Apache Spark to respond to the same problem - the problem of late data.

Continue Reading →

January 14, 2018 • Apache Beam

Late data in Apache Beam

Data, especially in streaming applications, can very often arrive on late to the processing pipeline. Despite of that, Apache Beam is able to handle this case pretty easily thanks to watermark mechanism.

Continue Reading →

late data articles

Apache Spark 2.4.0 features - watermark configuration

Apache Spark Structured Streaming and watermarks

Late data in Apache Beam