Forum Discussion
LTAP Direct Writes: Loading Terabytes into Lakebase Postgres Without Touching Your Live Application
A new beta capability that removes the classic bottleneck of bulk-loading data into a live Postgres database on Azure Databricks.
On October 8, 2026, Databricks published details on LTAP Direct Writes, a Beta capability of the LTAP (Lake Transactional/Analytical Processing) architecture that accelerates large data loads into Lakebase Postgres.
Until now, bulk-loading billions of rows into a live transactional Postgres database meant competing directly with the application for the same CPU and memory. Every write, regardless of its origin, has to pass through a single primary instance. One Databricks customer loading close to 1 billion rows per day through Synced Tables needed more than 8 hours per load, and had to overprovision OLTP compute just to absorb that peak without degrading the application.
LTAP Direct Writes attacks this problem by moving the heavy construction work out of the primary instance entirely. Databricks reports loading 1 TB of data in under 5 minutes with this approach, up to 147 times faster, without consuming resources from the live application.
[INSERT IMAGE HERE using the editor's own Insert Image button, not pasted markup. Upload the file, or if the editor accepts an external URL, use: https://www.databricks.com/sites/default/files/blog_images/load-tbs-in-minutes-into-postgres-blog-img-6.png . Suggested alt text: "LTAP architecture diagram showing Lakebase compute streaming the WAL to a storage layer that transcodes data into columnar Parquet, readable through Delta and Iceberg".]
1. Why a single Postgres primary cannot scale bulk loads linearly
In any standard Postgres deployment, every write, no matter where it originates, passes through the primary instance. This is not a configuration limitation, it is the engine's transactional guarantee: the primary has to coordinate commit order, maintain consistency, and manage the write-ahead log as a single sequence.
Throwing more parallelism at the client side does not help much when the bottleneck sits in the one process that can actually write. This is why bulk loads of billions of rows have historically competed directly with the production application for the same CPU and memory on the primary, forcing a choice between load throughput and application latency.
2. What LTAP Direct Writes changes
LTAP separates compute from storage in Lakebase. As data is materialized to object storage, it is transcoded from Postgres's row-oriented representation into columnar Parquet, readable through Delta or Iceberg. This creates a single logical copy of the data, served by two specialized engines: Postgres for transactional workloads, an analytical engine for OLAP, with no external CDC pipeline to keep in sync.
Synced Tables, the capability that serves lakehouse data back into Lakebase, is one concrete implementation of this architecture, and it is exactly where LTAP Direct Writes comes in to accelerate the initial load.
The mechanism moves data construction out of the primary and runs it in parallel, outside of it:
Parallel page building: Each Spark executor runs an isolated, sandboxed Postgres instance in binary-upgrade mode, the same mechanism `pg_upgrade` uses internally. Each sandbox reuses the destination's catalog OIDs, and the driver assigns non-overlapping OID ranges so concurrent objects cannot collide.
Frozen binary COPY: Inside each sandbox, a binary `COPY` with `FREEZE` builds that executor's slice of the heap. Frozen tuples are treated as already committed, so the sandbox does not need to maintain transaction history.
Per-page checksums: Each worker computes a checksum for its own constructed page. Lakebase's pageserver validates this on ingest, so a corrupted page never becomes authoritative.
Index building without downloading the full heap: Each worker exports key columns and tuple IDs to object storage. A custom table access method feeds these into Postgres's standard B-tree builder, shifting each tuple pointer by the cumulative size of earlier slices. The full heap never has to be downloaded back to build the index.
Atomic handoff: The Spark driver writes a manifest and calls a single SQL function. The primary records the import as a compact WAL entry, carrying only the import's description rather than the full data volume. The pageserver claims the files, warms them on local SSD, and the primary atomically swaps the staged data into the user-visible table. Until that commit, the imported table stays isolated.
The team behind this is explicit that the technique uses a Postgres extension and a custom table access method, without modifying Postgres core.
3. What this unlocks for Azure Databricks users specifically
For teams already running Lakebase on Azure Databricks, this closes a gap that used to force an uncomfortable choice: either accept a multi-hour load window with degraded application performance, or keep bulk analytical datasets out of Lakebase entirely and pay the cost of a separate reverse-ETL pipeline just to avoid that degradation.
With Direct Writes, feeding a Lakebase-backed application with a large batch of enriched lakehouse data (a nightly scoring table, a freshly computed feature set, a reference dataset rebuilt from scratch) no longer means scheduling that load for a maintenance window. It becomes a background operation that the live application does not feel.
4. Hands-on: creating a Synced Table with Direct Writes
The capability that benefits from Direct Writes today is the creation of a Synced Table. Direct Writes itself is currently a checkbox in the Create synced table dialog in the UI, not yet an explicit field documented in the CLI or SDK JSON spec, so the examples below show the general synced table creation flow that Direct Writes accelerates.
Creating the synced table with the Databricks CLI
databricks postgres create-synced-table main.sales.orders_sync \
--json '{
"spec": {
"source_table_full_name": "main.sales.orders",
"branch": "projects/my-project/branches/production",
"primary_key_columns": ["order_id"],
"scheduling_policy": "SNAPSHOT",
"postgres_database": "mydb",
"create_database_objects_if_missing": true
}
}'Creating the synced table with the Python SDK
from databricks.sdk import WorkspaceClient
from databricks.sdk.service.postgres import (
SyncedTable,
SyncedTableSyncedTableSpec,
SyncedTableSyncedTableSpecSyncedTableSchedulingPolicy,
)
w = WorkspaceClient()
synced_table = w.postgres.create_synced_table(
synced_table=SyncedTable(spec=SyncedTableSyncedTableSpec(
source_table_full_name="main.sales.orders",
branch="projects/my-project/branches/production",
primary_key_columns=["order_id"],
scheduling_policy=SyncedTableSyncedTableSpecSyncedTableSchedulingPolicy.SNAPSHOT,
postgres_database="mydb",
create_database_objects_if_missing=True,
)),
synced_table_id="main.sales.orders_sync",
).wait()
print(f"Synced table created: {synced_table.name}")Querying the loaded data directly from Postgres
Once the load completes, applications query the result with a standard Postgres driver, no different from any other table:
SELECT *
FROM sales.orders_sync
WHERE order_id = 48213;5. Checking load status programmatically
For a load that might take anywhere from seconds to minutes depending on volume, polling status through the SDK is more practical than watching the UI:
from databricks.sdk import WorkspaceClient
w = WorkspaceClient()
table = w.postgres.get_synced_table("synced_tables/main.sales.orders_sync")
print(f"State: {table.status.detailed_state}")
print(f"Last sync: {table.status.last_sync_time}")
print(f"Message: {table.status.message}")This is useful to wire into a Lakeflow Job as a dependency gate: downstream tasks can wait for `detailed_state` to report a completed sync before reading from the table, instead of assuming a fixed load duration.
6. What Direct Writes accelerates, and what it does not yet
LTAP Direct Writes accelerates the initial load in every sync mode. For subsequent syncs, the acceleration applies specifically to Snapshot mode, because that mode performs a full refresh on every cycle. In Triggered and Continuous modes, later updates are already incremental through Change Data Feed, not bulk loads, so Direct Writes only matters for their initial load.
The team behind the feature is upfront about its current scope:
- The 147x benchmark measures heap-building time only. Index building is not parallelized the same way yet, and the team describes this as active, ongoing work.
- LTAP Direct Writes is Beta, and requires a Lakebase project running Postgres 16, 17, or 18.
- It can only be selected when a synced table is created, not added to an existing one. To use it with an existing synced table, you delete it and create a new one.
- A table with many secondary indexes will likely not see the same proportional gain as the simple table used in the published benchmark, since index construction is the part still being parallelized.
7. My perspective on this change
The performance number is the headline, but the architectural detail that matters more is where the heavy work moved to. The load no longer competes with your live application inside the primary instance, it becomes isolated Spark work that happens alongside it. That is the problem the cited customer actually had: not "slow load" in isolation, but "slow load competing for resources with production."
I would test this first on a table without complex secondary indexes before assuming the 147x gain repeats identically on any schema, given the team's own disclosure that index building is still catching up to the heap-building speedup.
For Azure Databricks teams already committed to Lakebase as their operational Postgres, this is one more argument for keeping large, bulk-refreshed datasets inside Lakebase instead of routing them through a separate reverse-ETL tool just to avoid the old load-time penalty.
Conclusion
LTAP Direct Writes is not an isolated feature, it is separating compute from storage applied all the way down into the bulk-load process itself. For teams already using Synced Tables and feeling the weight of large loads on the primary, it is worth testing the Beta before assuming overprovisioning the OLTP side is the only way out.
Official references
Databricks Blog, "Load terabytes of data in minutes into Lakebase Postgres"
Microsoft Learn, "LTAP architecture - Azure Databricks"
Microsoft Learn, "Serve lakehouse data with synced tables"
Microsoft Learn, "Lakebase architecture - Azure Databricks"
Databricks Blog, "Lakebase LTAP: rethinking database storage"