Lakeflow Jobs, when it can and can't be a default data orchestration setup?

You have set up your Databricks platform and are about adding a new epic to your backlog, "Implement a data orchestration layer". Before you hit the add button, let's stop for a while and see if you really need this dedicated layer.

4-day workshop · In-person or online

What would it take for you to trust your Databricks pipelines in production?

A 3-day bug hunt on a 3-person team costs up to €7,200 in lost engineering time. This workshop teaches you to prevent that — unit tests, data tests, and integration tests for PySpark and Databricks Lakeflow, including Spark Declarative Pipelines.

Unit, data & integration tests
Medallion architecture & Lakeflow SDP
Max 10 participants · production-ready templates
See the full curriculum → €7,000 flat fee · cohort of up to 10
Bartosz Konieczny
Bartosz
Konieczny

💎 This blog post completes a great article by Daniel Beach Databricks Workflows vs Apache Airflow... or both?.

Lakeflow Jobs 101

If you needed to explain Lakeflow Jobs orchestration on the go, you could say:

Can Lakeflow Jobs be your default choice for data orchestration? As usual, the answer is it depends.

When it's enough

When you build your data platform on Databricks, Lakeflow Jobs should be enough. Generally, as long as you don't need to leave Databricks environment too much, leveraging Lakeflow Jobs will be the best choice, and so for several reasons:

When you need more

Even though Lakeflow Jobs orchestrator is a great fully-managed and well-integrated component on Databricks, there are use cases when it won't be enough. Typically, when you consider:

To answer the initial question if Lakeflow Jobs can be your single orchestrator, well, it depends! It depends on the vision of your final data platform; if it's fully Databricks, there is no sense of bringing another data orchestration component. If Databricks is only a part of your ecosystem, then an additional low-level orchestrator may be a better choice. Same for the complexity; if you can cover all your use cases pretty easily with Lakeflow Jobs capacities, no need to make your life harder with another data orchestrator. If defining your orchestration logic with Lakeflow Jobs takes three straight days every time because you need to write a lot of code, test, and package it, then devoting this time to bootstrap a dedicated orchestrator might be a better option - particularly now when it comes as a native cloud data offering.

Data Engineering Design Patterns

Looking for a book that defines and solves most common data engineering problems? I wrote one on that topic! You can read it online on the O'Reilly platform, or get a print copy on Amazon.

I also help solve your data engineering problems contact@waitingforcode.com đź“©