Articles about Spark SQL internals on waitingforcode.com

March 25, 2017 • Apache Spark SQL

Generated code in Spark SQL

One of powerful features of Spark SQL is dynamic generation of code. Several different layers are generated and this post explains some of them.

Continue Reading →

March 25, 2017 • Apache Spark SQL

Spark Project Tungsten

Even if Project Tungsten was started in Spark 1.5 and Spark's current version is 2.1 at the time of writing, it's good to know what precious this Project brought to Spark.

Continue Reading →

February 26, 2017 • Apache Spark SQL

Catalyst Optimizer in Spark SQL

The use of Dataset abstraction is not a single difference between structured and unstructured data processing in Spark. Apart of that, Spark SQL uses a technique helping to get results faster.

Continue Reading →

Spark SQL internals articles

Generated code in Spark SQL

Spark Project Tungsten

Catalyst Optimizer in Spark SQL