adls
2 TopicsBest Way to Exclude Delta _delta_log Folders from Purview Scans?
Hi Purview Community, We're scanning ADLS Gen2 assets with Microsoft Purview Data Map and have encountered issues with _delta_log folders being ingested as catalog assets. Initially, _delta_log folders were included in scans and were ingested into the Data Map. This appears to have introduced a large number of technical assets that are not useful for business users and may be contributing to unexpected metadata and row count discrepancies. Questions What is the recommended ignore pattern to exclude all _delta_log folders from ADLS Gen2 / Fabric scans? After updating a scan rule to exclude _delta_log, how can previously ingested _delta_log assets be completely removed from the Data Map? Is deleting the assets from Data Map sufficient? Is there a metadata purge process? Are there any known limitations where deleted _delta_log assets still affect resource sets, lineage, profiling, DQ, or row count calculations? Has anyone experienced inflated row counts or profiling results after _delta_log folders were previously scanned? For example, we're seeing significant differences between Purview-reported row counts and source system row counts even after excluding _delta_log folders and deleting the corresponding assets. Any guidance, best practices, or Microsoft documentation references would be greatly appreciated. Thanks!44Views0likes0CommentsLoading Parquet and Delta files into Azure Synapse using ADB or Azure Synapse?
I have a below case scenario. We are using Azure Databricks to pull data from several sources and generate the Parquet and Delta files and loaded them into our ADLS Gen2 Containers. We are now planning to create our data warehouse inside Azure Synapse SQL Pools, where we will create external tables for dimension tables which will use delta files and hash distributed fact tables using Parquet files. Now, the question is, to automate this data warehousing loading activity, which method is better? Is it better to use Azure Databricks to write our transformation logic to create dim and fact tables and load them regularly inside Azure Synapse SQL pools (or) is it better to use Azure Synapse to write our transformation logic to create dim and fact tables and load them regularly inside Azure Synapse SQL pools. Please help.854Views0likes1Comment