Forum Discussion

Wiliam_Rosa's avatar
Sep 30, 2026

Lakeflow in Azure Databricks

Anyone who has ever built a data engineering stack from scratch knows the pattern: one ingestion tool to pull data from CRM and ERP, a separate orchestrator to stitch together the dependencies between jobs, notebooks or scripts doing the actual transformation, and a fourth system just to check whether all of that ran correctly last night. Each piece comes from a different vendor, with its own authentication, its own logs, and its own SLA. When something breaks at three in the morning, the first job isn't fixing the problem, it's figuring out which of the four tools it lives in.

Lakeflow attacks that fragmentation head-on: it brings ingestion, transformation, and orchestration together into a single surface inside Azure Databricks, with Unity Catalog providing end-to-end governance. This isn't marketing repackaging of three products that already existed separately, it's the promise that the dependency between ingestion and transformation, column-level lineage, and access control stop being reconciled by hand across systems and start existing natively in the same graph.

The three pieces that make up Lakeflow

Lakeflow Connect covers ingestion, with point-and-click connectors for SaaS applications (Salesforce, Workday, ServiceNow), databases (SQL Server), and messaging, plus Zerobus Ingest, a serverless direct-write API for anyone with events coming from their own applications who doesn't want to stand up a message bus in the middle.

Spark Declarative Pipelines is the transformation layer: instead of hand-writing the orchestration of streaming tables and materialized views, you declare the target table and the business logic, and the engine infers the dependency graph on its own, builds the DAG, and takes care of operational details like backfill and versioning. The name has changed over time (anyone who has followed the platform for a while will recognize it as the direct evolution of what used to be Delta Live Tables), but the underlying mechanism, declaring the desired result instead of the procedural step-by-step, remains the central idea.

Lakeflow Jobs orchestrates all of it as a unified DAG: SQL workloads, Python code, declarative pipelines, dashboards, and external systems live in the same job definition, with data-driven triggers (a table was updated, a file landed in a volume), control-flow tasks, and no-code backfill runs.

Hands-on: a simple end-to-end pipeline

To make it concrete, here's how the three pieces fit together in a basic ingestion and declarative transformation pipeline:

import dlt
from pyspark.sql.functions import col

# Streaming table: incremental ingestion via Auto Loader
@dlt.table(
    comment="Raw order ingestion via Auto Loader"
)
def orders_bronze():
    return (
        spark.readStream.format("cloudFiles")
        .option("cloudFiles.format", "json")
        .load("/Volumes/sales/bronze/orders_raw")
    )

# Materialized view: validation and enrichment
@dlt.table(
    comment="Validated orders, with a quality filter"
)
@dlt.expect_or_drop("positive_amount", "total_amount > 0")
def orders_silver():
    return (
        dlt.read_stream("orders_bronze")
        .withColumn("total_amount", col("total_amount").cast("decimal(10,2)"))
        .filter(col("customer_id").isNotNull())
    )

This snippet doesn't run on its own, it becomes a task inside a Lakeflow Job that can also kick off an AI/BI dashboard as soon as the "orders_silver" table is updated, using a data-driven trigger instead of a fixed time-based schedule. That binding between ingestion, transformation, and the next step in the flow, all in the same place, is what eliminates much of the glue work that today lives in a separate Airflow or Azure Data Factory.

Governance without manual reconciliation

The point that usually goes unnoticed by anyone looking only at the product surface is what happens underneath with Unity Catalog. Because ingestion, transformation, and orchestration share the same catalog, data lineage is captured end to end automatically, from the source table in SQL Server all the way to the final dashboard, without relying on manual annotation or on a third-party lineage system trying to infer relationships from logs.

My take: this unification of governance is, in my view, Lakeflow's strongest argument, stronger even than the convenience of having a ready-made Workday connector. I've seen more than one data engineering project fail an audit not because the pipeline was wrong, but because nobody could prove with confidence where a specific number had come from, and how many transformations sat in between. Having that guaranteed structurally, rather than as a manual documentation process that is always out of date, changes the kind of conversation you have with the compliance team.

System Tables come in as the observability piece: instead of building your own monitoring dashboard by querying each separate tool's API, you can build alerts and health reports directly on top of native SQL tables, with retention and schema standardized by Databricks itself.

Cost: where the promise requires your own verification

Databricks publishes some impressive cost-reduction figures, including a case of up to 83% lower ETL cost and up to 90x performance improvement reported by a specific customer. Those numbers deserve the standard skepticism any vendor benchmark does: they usually reflect a specific migration, from a poorly optimized stack to a new, well-tuned configuration. In practice, I'd test the real savings by running the same representative workload from your own environment on old vs. new, with equivalent cluster policies and compute mode (Performance vs. Standard on serverless compute), before using the marketing number in a budget justification to leadership.

The mechanism behind the savings is real and makes technical sense: serverless compute with automatic optimization eliminates cold starts between sequential tasks in the same job, and granular per-task resource control avoids over-provisioning a cluster for a lightweight stage of the pipeline. That's different from guaranteeing that any migration will hit an 83% reduction.

What this doesn't solve

Consolidating ingestion, transformation, and orchestration into one platform doesn't remove the need for well-thought-out data modeling, nor does it replace architecture decisions about partitioning and clustering large tables. It also doesn't, on its own, solve the problem for teams that already have a heavy investment in Airflow with custom plugins that are hard to port, migrating legacy orchestration is still manual rewrite work, even with the dbt task and third-party integrations easing part of the path. And Lakeflow Connect's point-and-click connectors cover a specific set of sources, legacy systems or proprietary APIs without a native connector still depend on custom ingestion via Auto Loader or your own API.

Bottom line

Lakeflow doesn't invent a new data engineering paradigm, it removes the friction of coordinating three different tools to do what was always conceptually a single thing: bring data in from outside, transform it with confidence, and deliver it in the right shape to the next consumer. In practice: for anyone on Azure Databricks who already feels the pain of maintaining Data Factory, a third-party orchestrator, and a separate lineage catalog, it's worth piloting on a real pipeline before promising Databricks' cost-reduction number to anyone.

References

No RepliesBe the first to reply