azure
8172 TopicsComplete Guide to Azure Databricks Cost Optimization for Data Engineers
Co-Authored by: Aladdin Alchalabi AladdinAlchalabi, Sanjeev Nair Sanjeev Nair and Rafia Aqil Rafia_Aqil This guide walks through a proven approach to Azure Databricks cost optimization, structured in three phases: 1. Discovery, 2. Cluster/Data/Code Best Practices, and 3. Team Alignment & Next Steps. Phase 1: Discovery Assessing Your Current State The following questions are designed to guide your initial assessment and help you identify areas for improvement. Documenting answers to each will provide a baseline for optimization and inform the next phases of your cost management strategy. Environment & Organization Cluster Management Cost Optimization Data Management Performance Monitoring Future Planning What is the current scale of your Databricks environment? How many workspaces do you have? How are your workspaces organized (e.g., by environment type, region, use case)? How many clusters are deployed? How many users are active? What are the primary use cases for Databricks in your organization? Data engineering Data science Machine learning Business intelligence How are clusters currently managed? Manual configuration Automated scripts Databricks REST API Cluster policies What is the average cluster uptime? Hours per day Days per week What is the average cluster utilization rate? CPU usage Memory usage What is the current monthly spend on Databricks? Total cost Breakdown by workspace Breakdown by cluster What cost management tools are currently in use? Azure Cost Management Third-party tools Are there any existing cost optimization strategies in place? Reserved instances Spot instances Cluster auto-scaling What is the current data storage strategy? Data lake Data warehouse Hybrid What is the average data ingestion rate? GB per day Number of files What is the average data processing time? ETL jobs Machine learning models What types of data formats are used in your environment? Delta Lake Parquet JSON CSV Other formats relevant to your workloads What performance monitoring tools are currently in use? Databricks Ganglia Azure Monitor Third-party tools What are the key performance metrics tracked? Job execution time Cluster performance Data processing speed Are there any planned expansions or changes to the Databricks environment? New use cases Increased data volume Additional users What are the long-term goals for Databricks cost optimization? Reducing overall spend Improving resource utilization & cost attribution Enhancing performance Understanding Databricks Cost Structure Total Cost = Cloud Cost + DBU Cost Cloud Cost: Compute (VMs, networking, IP addresses), storage (ADLS, MLflow artifacts), other services (firewalls), cluster type (serverless compute, classic compute) DBU Cost: Workload size, cluster/warehouse size, photon acceleration, compute runtime, workspace tier, SKU type (Jobs, Delta Live Tables, All Purpose Clusters, Serverless), model serving, queries per second, model execution time Diagnose Cost and Issues Effectively diagnosing cost and performance issues in Databricks requires a structured approach. Use the following steps and metrics to gain visibility into your environment and uncover actionable insights. 1. Identify Costly Workloads Account Console Usage Reports: Review usage reports to identify usage breakdowns by product, SKU name, and custom tags. Usage Breakdown by Product and SKU: Helps you understand which services and compute types (clusters, SQL warehouses, serverless options) are consuming the most resources. Custom Tags for Attribution: Tags allow you to attribute costs to teams, projects, or departments, making it easier to identify high-cost areas. Workflow and Job Analysis: By correlating usage data with workflows and jobs, you can pinpoint long-running or resource-heavy workloads that drive costs. Focus on Long-Running Workloads: Examine workloads with extended runtimes or high resource utilization. Key Question: Which pipelines or workloads are driving the majority of your costs? Governance Hub: is a centralized, account-level UI for monitoring and managing governance across Databricks. Note this is a beta feature, we will update this article as we get more information and use-cases on this feature. Now That You’ve Identified Long-Running Workloads, Review These Key Areas: 2. Review Cluster Metrics CPU Utilization: Track guest, iowait, idle, irq, nice, softirq, steal, system, and user times to understand how compute resources are being used. Memory Utilization: Monitor used, free, buffer, and cached memory to identify over- or under-utilization. Key Question: Is your cluster over- or under-utilized? Are resources being wasted or stretched too thin? 3. Review SQL Warehouse Metrics Live Statistics: Monitor warehouse status, running/queued queries, and current cluster count. Time Scale Filter: Analyze query and cluster activity over different time frames (8 hours, 24 hours, 7 days, 14 days). Peak Query Count Chart: Identify periods of high concurrency. Completed Query Count Chart: Track throughput and query success/failure rates. Running Clusters Chart: Observe cluster allocation and recycling events. Query History Table: Filter and analyze queries by user, duration, status, and statement type. Key Question: Is your SQL Warehouse over- or under-utilized? Are resources being wasted or stretched too thin? 4. Review Spark UI Stages Tab: Look for skewed data, high input/output, and shuffle times. Uneven task durations may indicate data skew or inefficient data handling. Jobs Timeline: Identify long-running jobs or stages that consume excessive resources. Stage Analysis: Determine if stages are I/O bound or suffering from data skew/spill. Executor Metrics: Monitor memory usage, CPU utilization, and disk I/O. Frequent garbage collection or high memory usage may signal the need for better resource allocation. 4.1. Spark UI: Storage & Jobs Tab Storage Level: Check if data is stored in memory, on disk, or both. Size: Assess the size of cached data. Job Analysis: Investigate jobs that dominate the timeline or have unusually long durations. Look for gaps caused by complex execution plans, non-Spark code, driver overload, or cluster malfunction. 4.2. Spark UI: Executor Tab Storage Memory: Compare used vs. available memory. Task Time (Garbage Collection): Review long tasks and garbage collection times. Shuffle Read/Write: Measure data transferred between stages. 5. Additional Diagnostic Methods System Tables in Unity Catalog: Query system tables for cost attribution and resource usage trends. Cost Observability Queries Tagging Analysis: Use tags to identify which teams or projects consume the most resources. Dashboards & Alerts: Set up cost dashboards and budget alerts for proactive monitoring. Phase 2: Cluster/Code/Data Best Practices Alignment Cluster UI Configuration and Cost Attribution Effectively configuring clusters/workloads in Databricks is essential for balancing performance, scalability, and cost. Tunning settings and features when used strategically can help organizations maximize resource efficiency and minimize unnecessary spending. Key Configuration Strategies 1. Reduce Idle Time: Clusters to incur costs even when not actively processing workloads. To avoid paying for unused resources: Enable Auto-Terminate: Set clusters automatically shut down after a period of inactivity. This simple setting can significantly reduce wasted spending. Enable Autoscaling: Workloads fluctuate in size and complexity. Autoscaling allows clusters to dynamically adjust the number of nodes based on demand: Automatic Resource Adjustment: Scale up for heavy jobs and scale down for lighter loads, ensuring you only pay for what you use. It significantly enhances cost efficiency and overall performance. For serverless and streaming, using Delta Live Tables with autoscaling is recommended. This approach leads to better resource management and reliability. Use Spot Instances: For batch processing and non-critical workloads, spot instances offer substantial cost savings: Lower VM Costs: Spot instances are typically much cheaper than standard VMs. However, they are not recommended for jobs requiring constant uptime due to potential interruptions. Considerations: Azure Spot VMs are intended for non-critical, fault-tolerant tasks. They can be evicted without notice, risking production stability. No SLA guarantees mean potential downtime for critical applications. Using Spot VMs could lead to reliability issues in production environments. Leverage Photon Engine: Photon is Databricks’ high-performance, vectorized query engine: Accelerate Large Workloads: Photon can dramatically reduce runtime for compute-intensive tasks, improving both speed and cost efficiency. Keep Runtimes Up to Date: Using the latest Databricks runtime ensures optimal performance and security: Benefit from Improvements: Regular updates include performance enhancements, bug fixes, and new features. Apply Cluster Policies: Cluster policies help standardize configurations and enforce cost controls across teams: Governance and Consistency: Policies can restrict certain settings, enforce tagging, and ensure clusters are created with cost-effective defaults. Optimize Storage: type impacts both performance and cost: Switch from HDDs to SSDs: SSDs provide faster caching and shuffle operations, which can improve job efficiency and reduce runtime. Tag Clusters for Cost Attribution: Tagging clusters enables granular tracking and reporting: Visibility and Accountability: Use tags to attribute costs to specific teams, projects, or environments, supporting better budgeting and chargeback processes. Select the Right Cluster Type: Different workloads require different cluster types, see table below for Serverless vs Classic Compute: Feature Classic Compute Serverless Compute Control Full control over config & network Minimal control, fully managed by Databricks Startup Time Slower (unless pre-warmed) Instant Cost Model Hourly, supports reservations Pay-per-use, elastic scaling Security VNet injection, private endpoints NCC-based private connectivity Best For Heavy ETL, ML, compliance workloads Interactive queries, unpredictable demand Job Clusters: Ideal for scheduled jobs and Delta Live Tables. All-Purpose Clusters: Suited for ad-hoc analysis and collaborative work. Single-Node Clusters: Efficient for simple exploratory data analysis or pure Python tasks. Serverless Compute: Scalable, managed workloads with automatic resource management. 11. Monitor and Adjust Regularly: review cluster metrics and query history: Continuous Optimization: Use built-in dashboards to monitor usage, identify bottlenecks, and adjust cluster size or configuration as needed. Code Best Practices Avoid Reprocessing Large Tables Use a CDC (Change Data Capture) architecture with Delta Live Tables (DLT) to process only new or changed data, minimizing unnecessary computation. Ensure Code Parallelizes Well Write Spark code that leverages parallel processing. Avoid loops, deeply nested structures, and inefficient user-defined functions (UDFs) that can hinder scalability. Reduce Memory Consumption Tweak Spark configurations to minimize memory overhead. Clean out legacy or unnecessary settings that may have carried over from previous Spark versions. Prefer SQL Over Complex Python Use SQL (declarative language) for Spark jobs whenever possible. SQL queries are typically more efficient and easier to optimize than complex Python logic. Modularize Notebooks Use %run to split large notebooks into smaller, reusable modules. This improves maintainability. Use LIMIT in Exploratory Queries When exploring data, always use the LIMIT clause to avoid scanning large datasets unnecessarily. Monitor Job Performance Regularly review Spark UI to detect inefficiencies such as high shuffle, input, or output. Review the below table for optimization opportunities: Spark stage high I/O - Azure Databricks | Microsoft Learn Databricks Code Performance Enhancements & Data Engineering Best Practices By enabling the below features and applying best practices, you can significantly lower costs, accelerate job execution, and build Databricks pipelines that are both scalable and highly reliable. For more guidance review: Comprehensive Guide to Optimize Data Workloads | Databricks. Feature / Technique Purpose / Benefit How to Use / Enable / Key Notes Disk Caching Accelerates repeated reads of Parquet files Set spark.databricks.io.cache.enabled = true Dynamic File Pruning (DFP) Skips irrelevant data files during queries, improves query performance Enabled by default in Databricks Low Shuffle Merge Reduces data rewriting during MERGE operations, less need to recalculate ZORDER Use Databricks runtime with feature enabled Adaptive Query Execution (AQE) Dynamically optimizes query plans based on runtime statistics Available in Spark 3.0+, enabled by default Deletion Vectors Efficient row removal/change without rewriting entire Parquet file Enable in workspace settings, use with Delta Lake Materialized Views Faster BI queries, reduced compute for frequently accessed data Create in Databricks SQL Optimize Compacts Delta Lake files, improves query performance Run regularly, combine with ZORDER on high-cardinality columns ZORDER Physically sorts/co-locates data by chosen columns for faster queries Use with OPTIMIZE, select columns frequently used in filters/joins Auto Optimize Automatically compacts small files during writes Enable optimizeWrite and autoCompact table properties Liquid Clustering Simplifies data layout, replaces partitioning/ZORDER, flexible clustering keys Recommended for new Delta tables, enables easy redefinition of clustering keys File Size Tuning Achieve optimal file size for performance and cost Set delta.targetFileSize table property Broadcast Hash Join Optimizes joins by broadcasting smaller tables Adjust spark.sql.autoBroadcastJoinThreshold and spark.databricks.adaptive.autoBroadcastJoinThreshold Shuffle Hash Join Faster join alternative to sort-merge join Prefer over sort-merge join when broadcasting isn’t possible, Photon engine can help Cost-Based Optimizer (CBO) Improves query plans for complex joins Enabled by default, collect column/table statistics with ANALYZE TABLE Data Spilling & Skew Handles uneven data distribution and excessive shuffle Use AQE, set spark.sql.shuffle.partitions=auto, optimize partitioning Data Explosion Management Controls partition sizes after transformations (e.g., explode, join) Adjust spark.sql.files.maxPartitionBytes, use repartition() after reads Delta Merge Efficient upserts and CDC (Change Data Capture) Use MERGE operation in Delta Lake, combine with CDC architecture Data Purging (Vacuum) Removes stale data files, maintains storage efficiency Run VACUUM regularly based on transaction frequency Phase 3: Team Alignment and Next Steps Implementing Cost Observability and Taking Action Effective cost management in Databricks goes beyond configuration and code—it requires robust observability, granular tracking, and proactive measures. Below outlines how your teams can achieve this using system tables, tagging, dashboards, and actionable scripts. Cost Observability with System Tables Databricks Unity Catalog provides system tables that store operational data for your account. These tables enable historical cost observability and empower FinOps teams to analyze spend independently. System Tables Location: Found inside the Unity Catalog under the “system” schema. Key Benefits: Structured data for querying, historical analysis, and cost attribution. Action: Assign permissions to FinOps teams so they can access and analyze dedicated cost tables. Enable Tags for Granular Tracking Tagging is a powerful feature for tracking, reporting, and budgeting at a granular level. Classic Compute: Manually add key/value pairs when creating clusters, jobs, SQL Warehouses, or Model Serving endpoints. Use cluster policies to enforce custom tags. Serverless Compute: Create budget policies and assign permissions to teams or members for serverless workloads. Action: Tag all compute resources to enable detailed cost attribution and reporting. Track Costs with Dashboards and Alerts Databricks offers prebuilt dashboards and queries for cost forecasting and usage analysis. Dashboards: Visualize spend, usage trends, and forecast future costs. Prebuilt Queries: Use top queries with system tables to answer meaningful cost questions. Budget Alerts: Set up alerts in the Account Console (Usage > Budget) to receive notifications when spend approaches defined thresholds. Build Culture of Efficiency To go beyond technical fixes and build a culture of efficiency, by focusing on the below strategic actions: Collaborate with Internal Engineers: Spend time with engineering teams to understand workload patterns and optimization opportunities. Peer Reviews and Code Audits: Conduct regular code review sessions and peer reviews to ensure best practices are followed for Spark jobs, data pipelines, and cluster configurations. Create Internal Best Practice Documentation: Develop clear guidelines for writing optimized code, managing data, and maintaining clusters. Make these resources easily accessible for all teams. Implement Observability Dashboards: Use Databricks’ built-in features to create dashboards that track spend, monitor resource utilization, and highlight anomalies. Set Alerts and Budgets: Configure alerts for long-running workloads and establish budgets using prebuilt Databricks capabilities to prevent cost overruns. 5. Azure Reservations and Azure Savings Plan When optimizing Databricks costs on Azure, it’s important to understand the two main commitment-based savings options: Azure Reservations and Azure Savings Plans. Both can help you reduce compute costs, but they differ in flexibility and how savings are applied. Which Should You Choose? Reservations are ideal if you have stable, predictable Databricks workloads and want maximum savings. Savings Plans are better if you expect your compute needs to change, or if you want a simpler, more flexible way to save across multiple services. Pro Tip: You can combine both options—use Reservations for your baseline, always-on Databricks clusters, and Savings Plans for bursty, variable, or new workloads. Summary Table: Action Steps It’s critical to monitor costs continuously and align your teams with established best practices, while scheduling regular code review sessions to ensure efficiency and consistency. Area Best Practice / Action System Tables Use for historical cost analysis and attribution Tagging Apply to all compute resources for granular tracking Dashboards Visualize spend, usage, and forecasts Alerts Set budget alerts for proactive cost management Scripts/Queries Build custom analysis tools for deep insights Cluster/Data/Code Review & Align Regularly review best practices, share findings, and align teams on optimization Save on your Usage Consider Azure Reservations and Azure Savings Plan705Views2likes0CommentsEnabling the Compliance Security Profile (CSP) for HIPAA on Azure Databricks
Microsoft Architect's: Aladdin Alchalabi AladdinAlchalabi, Kiran Raja KiranRaja, Peter Lenges PeterLenges, Jessica Reece jareece, Benjamin Coughtry bcoughtry, Anishek Kamal anishekkamal, Tayo Akigbogun takigbogun, Eric Kwashie ekwashie, Peter Lo PeterLo and Rafia Aqil Rafia_Aqil Peer Reviewed: Ted Kim tedkim and Arvind Periyasamy ArvindPeriyasamy Purpose and WHY Azure Databricks has put in place controls to meet the unique compliance needs of highly regulated industries. The requirement for the compliance security profile (CSP) is a joint effort between Microsoft and Databricks for Azure-Databricks workspaces. The value proposition of the compliance security profile is that it provides Customers significantly more hardening and security features. Mandatory Deadline: The Compliance Security Profile (CSP) becomes mandatory for processing HIPAA, HITRUST, and IRAP regulated data on Azure-Databricks by September 1, 2026. Key Dates: Enable the Compliance Security Profile and select HIPAA by September 1, 2026. Prepare for Azure Virtual Network encryption enforcement beginning February 1, 2027. Compliance Responsibility: Enabling CSP supports applicable technical controls but does not, by itself, establish HIPAA compliance. Compliance is a shared responsibility among the customer, Microsoft, and Databricks. Customers must evaluate their administrative, physical, and technical safeguards and confirm that an applicable Microsoft Business Associate Agreement is in place. Enabling CSP on Workspaces These requirements are checked and enforced on new workspaces today, with enforcement on existing workspaces expected in the future; where prerequisites are missing, clusters may fail to start. Prerequisite Requirement Costs There is a 10% cost of the Azure Databricks product spend within each workspace where CSP is enabled. **Review with your account team for any grace period during which the Enhanced Security & Compliance (ESC) add-on is available at no charge. After the grace period ends, a 10% DBU upcharge applies. Enhanced Security & Compliance add-on For existing workspaces: From Azure portal, click the Settings > Security & compliance on an existing Azure Databricks workspace: **Review Note #2 below Azure VNet encryption Azure Virtual Network encryption must be enabled on the Azure Databricks workspace VNet. Infrastructure as code: Update the encryption block on your VNet resource. In Terraform, that's azurerm_virtual_network. Azure portal: Toggle encryption on the VNet (Overview → Properties → Encryption). Command line: Enable it with the Azure CLI or PowerShell. **Review Note #4 below Supported VM instance types Use a VM series that supports VNet encryption and verify compatibility before enabling the profile. **This does not apply to serverless compute. NOTE: Confirm your workspace is using Premium Pricing tier. The profile can be enabled when a workspace is created or on an existing workspace, through the Azure portal, the Azure CLI, PowerShell, an ARM template, or Terraform. Only the Public Preview, Private Preview, and Beta features listed in this section are supported for workspaces with the compliance security profile enabled: Compliance security profile - Azure Databricks | Microsoft Learn Currently, the compliance security profile checks and enforces only the use of specific VM instance types, not the enablement of Azure Virtual Network encryption. Enforcement of the Azure Virtual Network encryption requirement begins on February 1, 2027, including on workspaces that already have the compliance security profile enabled. This flexibility shall allow customers more time to set up VNET encryption. This has been updated in documentation today (See ‘Important’ box). Regarding rollback, CSP can be reversed via a support ticket, if no regulated data has been processed on a particular workspace. Plan for possible effects on cluster startup, networking, feature availability, maintenance operations, and cost. A closer look at VNet encryption CSP is enabled per Databricks workspace, but VNet encryption is applied at the VNet level. Enabling it for a Databricks workload therefore affects every resource within that VNet, not just the workspace. A common approach in a hub-and-spoke design is to leave the hub VNet unencrypted and encrypt only the spoke VNet. The hub typically holds shared services such as the DNS resolver, while the spoke hosts the Databricks workspaces that require CSP. What does it mean for a VNet to be encrypted? An encrypted VNet is a security measure that protects VM-to-VM traffic. Data is encrypted in transit through a DTLS tunnel. This is platform-level encryption, applied automatically to traffic within your VNet and across peered VNets. It requires no changes to your operating system or applications. What happens to my VM-to-VM traffic? Qualifying VM-to-VM traffic is encrypted. Traffic involving unqualified instances simply keeps flowing unencrypted. The only enforcement available today is AllowUnencrypted. Important clarifications Encrypting the VNet does not guarantee all traffic within it will be encrypted. The only traffic that gets encrypted is VM-to-VM traffic where both the source and destination VMs are (1) on a supported SKU and (2) have Accelerated Networking enabled on the network interface. Encrypting the VNet does not drop or break traffic from unsupported SKUs. The only supported GA setting today is to allow unencrypted traffic, so non-qualifying traffic is still permitted; it just isn't encrypted. A future DropUnencrypted setting will drop that traffic instead for further hardening. It isn't available yet, and it's currently unknown whether it will become a required setting for CSP. Review the following recommended steps The steps below represent a validated implementation pattern. The exact network design can vary by environment, but the same prerequisite, isolation, and end-to-end validation principles should be applied. Validated implementation step Recommended approach and expected outcome Isolated sandbox workspace Enable CSP first in a representative non-production workspace. This avoids irreversible changes to DEV or production while the network topology, dependencies, VM compatibility, and operational behavior are validated. Enable CSP and select HIPAA Enable the Compliance Security Profile and select HIPAA under Settings > Security & compliance before processing PHI after September 1, 2026. Enable VNet encryption Enable VNet encryption. **Review Azure Virtual Network encryption limitations: What is Azure Virtual Network encryption? - Azure Virtual Network | Microsoft Learn Start a classic cluster Confirm that a classic cluster starts successfully after CSP and VNet encryption prerequisites are applied. This validates that the selected compute path and VM types remain operational. Validate storage connectivity Confirm storage connectivity continue to work. Confirm rollout readiness Proceed to DEV and production only after the complete private connectivity path, cluster startup, storage access, DNS resolution, data pipelines, and performance have been validated from end to end. Things to Review Enablement is permanent Enabling the compliance security profile, or adding a compliance standard, is intended to be a permanent change. You cannot remove the profile or an individual standard from a workspace that has ever processed regulated data; to revert, you must delete the workspace and create a new one. Validate the configuration in an isolated, representative non-production workspace before enabling DEV or production. Inventory and Assessment Identify Regulated Workspaces: Catalogue all existing Azure-Databricks workspaces. Determine which ones currently process, or are planned to process, data subject to HIPAA, HITRUST, or IRAP. Review Data Pipelines: Map out all data ingress and egress points for these identified workspaces, including connections to on-premises data sources, other cloud services, and external APIs. This helps identify potential network impacts. Verify Prerequisites Before Rollout: Confirm that selected VM instance types support VNet encryption and that every required CSP and networking setting is in place, because missing prerequisites can prevent clusters from starting. Enablement Method: Choose the appropriate tooling for enablement of Azure Portal, Azure CLI, PowerShell, ARM templates, or Terraform to ensure consistency and automation. Keep sensitive data out of customer-defined fields You are solely responsible for ensuring that PHI or other sensitive information is never entered into customer-defined input fields. These include workspace names, compute and resource names, tags, job names, job run names, network names, credential names, storage account names, and Git repository IDs or URLs, all of which may be stored, processed, or accessed outside the compliance boundary. What Changes After Enabling Compliance Security Profile On CSP-enabled workspaces, Partner-powered AI features are disabled by default and some assistive features such as Genie Code are also disabled; a workspace admin can re-enable them if required. In addition, only the specific preview features listed in the compliance security profile documentation are supported. No other Public Preview, Private Preview, or Beta feature may be used to process regulated data. Compliance Security Profile (CSP) enhances the security posture of Azure Databricks by enabling a hardened compute image, enhanced security monitoring, and automatic cluster updates. With automatic cluster updates enabled, classic compute resources are periodically updated and may restart during configured maintenance windows, so production schedules should be planned accordingly. Enhanced security monitoring deploys security monitoring agents on supported compute resources and generates logs that security teams can ingest and analyze. When deploying through ARM templates, CSP, enhancedSecurityMonitoring, and automaticClusterUpdate are configurable security and compliance settings that can be specified as part of the workspace deployment. Why GPU-accelerated compute may be affected Azure Databricks supports several GPU families, but Azure VNet encryption currently documents a narrower GPU list. Please review the list here: What is Azure Virtual Network encryption? - Azure Virtual Network | Microsoft Learn What happens if a customer does not enable Azure Databricks CSP? Azure Databricks has implemented additional controls to support the security and compliance requirements of highly regulated industries. CSP should not be characterized solely as a Microsoft initiative; it is part of the Azure Databricks security and compliance offering delivered by Microsoft and Databricks. CSP provides additional platform hardening and security capabilities, including a CIS Level 1 hardened compute image, automatic cluster updates, enhanced security monitoring, and TLS 1.2 or higher for relevant communications. Beginning September 1, 2026, CSP is required for Azure Databricks workspaces processing data subject to applicable standards, including HIPAA, HITRUST, and IRAP. If a customer chooses not to enable CSP while processing data subject to one of these standards, the workspace would not meet the documented Azure Databricks configuration requirements for that regulated workload. This does not, by itself, determine the customer's overall legal or regulatory compliance; customers should assess their obligations with their legal, compliance, and audit teams. References Compliance security profile: https://learn.microsoft.com/en-us/azure/databricks/security/privacy/security-profile Configure enhanced security and compliance settings: https://learn.microsoft.com/en-us/azure/databricks/security/privacy/enhanced-security-compliance HIPAA, Azure Databricks, Microsoft Learn: https://learn.microsoft.com/en-us/azure/databricks/security/privacy/hipaa What is Azure Virtual Network encryption: https://learn.microsoft.com/en-us/azure/virtual-network/virtual-network-encryption-overview Create a Virtual Network with encryption: https://learn.microsoft.com/en-us/azure/virtual-network/how-to-create-encryption?tabs Hashicorp azurerm_virtual_network: azurerm_virtual_network | Resources | hashicorp/azurerm | Terraform | Terraform Registry1.8KViews2likes0CommentsGetting first customers after publishing what actually worked for you?
I published two transactable SaaS offers earlier this month (Microsoft 365 security and compliance assessment). The technical side went fine, with good support from the ISV team. Getting the offers in front of customers is where I'm stuck. So far I've reached Rewards Tier 3 and activated the listing optimisation benefit. I understand customer references are the gate to co-sell readiness, which feels like a chicken-and-egg problem when you have none yet. For publishers further along than me: what actually produced your first few customers? Was it private offers to targeted tenants, co-sell once you had references, the Marketplace Rewards activities, or something outside the marketplace entirely? I'd rather spend my limited time on what works than keep guessing.12Views0likes0CommentsStreaming and Batch Data Architectures with Microsoft Fabric to Azure Databricks
Author's: Aladdin Alchalabi AladdinAlchalabi, Oscar Alvarado oscaralvarado and Rafia Aqil Rafia_Aqil Note: This article describes a solution idea. Your cloud architect can use this guidance to help visualize the major components for a typical implementation. Use this article as a starting point to design a well-architected solution that aligns with your workload’s specific requirements. As organizations adopt Microsoft Fabric as their unified analytics platform, it has become a leading path for ingesting both streaming and batch data into Azure Databricks. This article covers integration approaches -via Microsoft Fabric- and details the five Fabric-specific paths that connect OneLake/ADLS and Databricks for end-to-end data processing. Medallion Architecture The following data flow corresponds to the architecture diagram: Data is ingested through Microsoft Fabric (via Mirroring, RTI, or Data Factory) lands data into OneLake/ADLS. With the medallion pattern, consisting of Bronze, Silver, and Gold storage layers, organizations have flexible access and extendable data processing: Bronze – Raw data entry point. Data arrives in its source format and is converted to the open, transactional Delta Lake format. Silver – Optimized for BI and data science. ETL and stream processing tasks filter, clean, transform, join, and aggregate Bronze data into curated datasets using SQL, Python, R, or Scala. Gold – Enriched data ready for analytics and reporting. Analysts use Power BI, PySpark, SQL, or Excel for insights and queries. Fabric Integration Paths Note: This architecture establishes a complete loop-back between Microsoft Fabric and Azure Databricks, enabling Gold layer tables to be seamlessly mirrored back to Microsoft Fabric for dashboarding through Azure Databricks Mirroring. The following five paths connect Microsoft Fabric to Azure Databricks: Fabric Mirroring to OneLake – A low-cost, low-latency turnkey solution that creates a replica of data from operational sources (SQL Server, Azure Cosmos DB, Oracle) in OneLake. Handles the initial load and ongoing CDC changes automatically, keeping data continuously up to date. Fabric RTI to OneLake – Fabric Real-Time Intelligence ingests streaming event data into OneLake with sub-second latency, enabling real-time analytics on live event streams. Fabric Data Factory to OneLake – Orchestrates ingestion from diverse sources not covered by Mirroring (such as Sybase or REST APIs) and lands data in OneLake, ensuring complete source coverage. OneLake to Azure Databricks – Unity Catalog connections to OneLake, secured via Managed Identities from Microsoft Entra ID, allow Databricks to query OneLake data items as a native catalog without data duplication. Fabric Data Factory to Azure Databricks (direct) – Orchestrates ingestion from diverse sources directly into Azure Data Lake Storage (ADLS), where Azure Databricks picks up the data for medallion architecture processing. Design Considerations Area Updated guidance Direct RTI-to-Databricks integration There is still no broad GA direct integration where Fabric RTI and Databricks operate as one native real-time runtime. Integration should be positioned through open protocols, Event Hubs/Kafka-style patterns, OneLake, Delta, and federation. OneLake federation in Azure Databricks OneLake federation in Azure Databricks is now the key integration story. It allows Databricks Unity Catalog to query Fabric Lakehouse and Warehouse data in OneLake without copying it. Access is read-only and depends on Fabric tenant settings, workspace permissions, and Databricks Unity Catalog setup. RTI data availability to Databricks Data ingested through Fabric RTI can be made available to Databricks by landing or exposing the data into OneLake-backed items, especially Lakehouse/Warehouse patterns. Eventhouse data can be made available in OneLake in Delta format through OneLake availability, but Databricks OneLake federation should be validated against the specific Fabric item type and access path. Existing Databricks customers Existing Databricks customers do not need to abandon Databricks. They can use Fabric RTI as the event ingestion, real-time detection, operational alerting, and business action layer, while continuing to use Databricks for engineering, ML, advanced analytics, and Unity Catalog-governed access. Activator and business action Fabric Activator is the cleanest business-user action layer. It can monitor streaming events and trigger Teams messages, email, Power Automate flows, Fabric pipelines, notebooks, Spark jobs, Dataflows, UDFs, and other downstream actions. This is a strong differentiator because it lets business users act on events without waiting for batch analytics. Operations Agents Operations Agents are in preview and should be positioned carefully. They monitor real-time data from Eventhouse or ontology sources, surface insights, recommend actions, and can connect to Activator/Power Automate action paths. They are not simply a pre-ingestion decision engine before data lands anywhere; they work from configured Fabric knowledge/data sources. Before landing in Lakehouse For decisioning before Lakehouse persistence, use Eventstream processing and Activator rules on streams. For AI-assisted operational recommendations, use Operations Agents once the relevant data is available in Eventhouse or ontology. Requirement-Specific Notes Data Ingestion Microsoft Fabric Mirroring currently supports SQL Server, Azure Cosmos DB, and Oracle as source systems. For sources not yet supported by Mirroring—such as Sybase or REST APIs—use Fabric Data Factory pipelines to ensure full coverage across all data systems. Once data is in the landing zone with the correct format, Mirroring’s CDC replication starts automatically and manages the complexity of merging changes (updates, inserts, and deletes) into Delta tables, keeping data in Fabric continuously up to date. Learn more about open mirroring Storage Format and Time Travel OneLake supports Delta tables, enabling schema evolution and time travel across all data stored in the lakehouse. Learn more about OneLake and Delta tables Security Encryption at rest: OneLake automatically encrypts all data at rest using Microsoft-managed keys, compliant with FIPS 140-2 standards. Learn more Encryption in transit: All data in transit is encrypted using TLS 1.2 or higher, securing data movement between Fabric, OneLake, and Azure Databricks. Learn more Data Governance OneLake can be registered and scanned by Microsoft Purview, enabling cataloging of stored metadata and data quality profiling. This protects sensitive information, including PHI and PII, across ingestion and analytics workflows. Learn more about Purview with Fabric Lakehouse Operations and Monitoring Use the Fabric monitor hub to track pipeline health, Spark application performance, and ingestion job status across all Fabric workloads. Learn more about the Fabric monitor hub Scenario Details This architecture applies to any organization that needs to unify streaming and batch data at scale. Common characteristics include: Multiple operational data sources (databases, SaaS applications, event streams) A requirement to process both real-time and historical data in the same platform Governance and compliance requirements for sensitive data (PHI, PII, financial records) Analytics consumers spanning BI (Power BI), data science (Databricks notebooks), and ML workloads Potential Use Cases Healthcare and life sciences – PHI/PII protection via Purview; real-time patient telemetry + batch EHR analytics Financial services – Real-time fraud detection streams + batch regulatory reporting Retail and e-commerce – Streaming clickstream analytics + batch inventory and supply chain processing Energy and utilities – IoT sensor telemetry streaming + batch consumption analytics Next Steps Get started with Microsoft Fabric Mirroring Build an ETL pipeline with Lakeflow Declarative Pipelines Configure Unity Catalog with OneLake shortcuts Monitor Fabric pipelines with the Fabric monitor hub803Views2likes0CommentsSyksyn 2026 Tekniset ja myynnin Kumppanitunnit
Microsoftin tekniset ja myynnille suunnatut Kumppanitunnit järjestetään nykyään Microsoftin globaalilla Skilling Hub -sivustolla, josta ne ovat kätevästi saatavilla myöhemmin tallenteina materiaaleineen. Rekisteröidy Skilling Hub -portaaliin, josta löydät kaikki Microsoftin kumppanikoulutukset yhdessä paikassa eri kielillä tai tekstitettyinä. Rekisteröityessäsi voi valita ne kielet, kuten suomi, jollaista sisältöä haluat ensisijaisesti nähdä englannin kielisen koulutussisällön lisäksi. Kumppanitunti on joka toinen perjantai klo 10–11 järjestettävä Microsoftin kumppaniwebinaari, joka on tarkoitettu kaikille Microsoftin kumppaneille. Tekniset ja kaupalliset aiheet vuorottelevat ja olet tervetullut molempiin webinaareihin. Webinaareissa keskitymme Microsoftin ratkaisualueiden teknologioiden mielenkiintoisiin uutuuksiin, MAICPP-kumppaniohjelmaan, kumppanietuihin ja ratkaisumyyntiin. Microsoftin suomalaiset arkkitehdit, tuotepäälliköt, ratkaisumyyjät ja kumppanivastaavat ovat poimineet kiinnostavia ja hyödyllisiä aiheita, joita he vuorollaan esittelevät. Syksyn 2026 ohjelma Alla ovat suunnitellut päivät ja teemat Syksylle 2026, joiden tarkka aihe päivitetään aina lähempänä esityspäivää tälle sivulle. 18.9. Kaupallinen Kumppanitunti: Partner START FY27 - webinaari Rekisteröidy mukaan tästä linkistä: Etkö päässyt mukaan START FY27 Kickoff -tapahtumaan? Webinaarissa saat kattavan katsauksen tapahtumassa käsiteltyihin FY27: n liiketoimintamahdollisuuksiin, strategisiin painopisteisiin sekä Microsoftin kumppaneille suunnattuihin ohjelmiin ja resursseihin. Puhujat: Kalle Saarikannas, Liiketoimintajohtaja, AI Business Solutions Teemu Lainiola, Liiketoimintajohtaja, Cloud & AI Platforms Mereta Laukkanen, Sr Partner Development Manager Jonna Kaarlenkaski, Partner Development Manager Intern Jonna Fred-Jokela, Sr Solution Engineer Manager 2.10. Tekninen Kumppanitunti: Copilot Cowork: Copilot Credits ja kustannusten hallinta Rekisteröidy mukaan tästä linkistä! Tässä teknisessä kumppanitunnissa käymme läpi Copilot Coworkin toimintamallin, kulutuksen seurannan sekä kustannusten hallinnan. Lisäksi tarkastelemme, millaisia mahdollisuuksia kulutuspohjainen AI luo kumppaneiden palveluliiketoiminnalle. Puhujat: Henri Nevalainen, Microsoft Niko Hiltunen, IAMCP 9.10. Kaupallinen Kumppanitunti: Onko julkisen pilven suvereniteetti riittävä huomioiden sen kaikki hyödyt? Rekisteröidy mukaan tästä linkistä! Euroopan unioni valmistelee uusia pilvi- ja tekoälyinfrastruktuuria koskevia linjauksia. Samanaikaisesti organisaatiot pohtivat, miten digitaalinen suvereniteetti, tekoälyn käyttöönotto, kyberturvallisuus ja sääntelyvaatimukset voidaan sovittaa yhteen käytännössä. Tervetuloa ajankohtaiseen webinaariin, jossa tarkastelemme Euroopan muuttuvaa pilviympäristöä ja mitä digitaalinen suvereniteetti tarkoittaa suomalaisille organisaatioille käytännössä. Lisäksi kerromme viimeisimmät tilannetiedot pilvipalveluiden kvanttiturvallisuuden ja Suomen datakeskushankkeiden osalta. Puhujat: Juha Karppinen, National Technology Officer, Microsoft Timo Salminen, Partner Solution Architect, Microsoft Niko Hiltunen, IAMCP 22.10. Kaupallinen Kumppanitunti: Frontier Transformation - Next level IQ & Uusi näkökulma kumppanin rooliin AI-aikakaudella Rekisteröitymislinkki päivittyy tähän AI-transformaatio on siirtymässä yksittäisten työkalujen käyttöönotosta kohti laajempaa työn ja toimintamallien uudistamista. Kumppanitunnilla havainnollistamme käytännön esimerkin avulla, kuinka organisaation toimintaa voidaan tarkastella digitaalisen genomin kautta, organisaatiosta ja tiedosta aina työhön, mikrotehtäviin ja agentteihin asti. Samalla tarkastelemme, mitä muutos tarkoittaa kumppaneille: miten tunnistaa uusia asiakasarvon mahdollisuuksia, rakentaa vaikuttavampia AI-keskusteluja ja auttaa asiakkaita etenemään kohti agenttipohjaista toimintamallia. Tule mukaan katsomaan demo ja pohtimaan, mitä seuraavan vaiheen AI-transformaatio mahdollistaa sinun asiakkaillesi. Puhujat: Hans Dolk, Asiakkuusjohtaja / Defence & Intelligence, Public Safety and National Security Juha Liljeblom,Teknologia Strategisti / Defence & Intelligence, Public Safety and National Security 23.10. Kaupallinen Kumppanitunti: Uudistunut Microsoft Copilot: ominaisuudet, käyttötavat ja kustannusten hallinta Rekisteröitymislinkki päivittyy tähän Microsoft Copilot uudistuu merkittävästi. Copilotista rakentuu paikka, jossa käyttäjä voi kysyä, delegoida, rakentaa ja automatisoida työtä. Uudet Home-, Code- ja Autopilot-kokemukset laajentavat Copilotin käyttöä päivittäisestä AI-avustamisesta kohti pidempikestoista agenttityötä ja käyttäjien itse rakentamia ratkaisuja. Webinaarin tavoitteena on antaa kumppaneille selkeä kokonaiskuva siitä, miten uudistunut Copilot muuttaa asiakkaiden AI:n käyttötapoja, lisensointi- ja kulutuskeskustelua sekä kumppaneille avautuvia mahdollisuuksia. Puhujat: Harri Mikkanen, Microsoft Tuulia Piehl, Microsoft Henri Nevalainen, Microsoft 30.10. Tekninen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 6.11. Tekninen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 20.11. Kaupallinen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 4.12. Kaupallinen Kumppanitunti: Rekisteröitymislinkki päivittyy tähän Puhujat: 18.12. Tekninen Kumppanitunti: Agentit Azuressa Rekisteröitymislinkki päivittyy tähän Azure Copilot pitää sisällään uusia palveluiden elinkaarenhallintaan liittyviä agentteja. Tule kuulemaan, miten voit hyödyntää näitä omissa palveluissasi! Puhujat: Timo Salminen, Partner Solution Architect, Microsoft543Views0likes0CommentsLakeflow in Azure Databricks
Anyone who has ever built a data engineering stack from scratch knows the pattern: one ingestion tool to pull data from CRM and ERP, a separate orchestrator to stitch together the dependencies between jobs, notebooks or scripts doing the actual transformation, and a fourth system just to check whether all of that ran correctly last night. Each piece comes from a different vendor, with its own authentication, its own logs, and its own SLA. When something breaks at three in the morning, the first job isn't fixing the problem, it's figuring out which of the four tools it lives in. Lakeflow attacks that fragmentation head-on: it brings ingestion, transformation, and orchestration together into a single surface inside Azure Databricks, with Unity Catalog providing end-to-end governance. This isn't marketing repackaging of three products that already existed separately, it's the promise that the dependency between ingestion and transformation, column-level lineage, and access control stop being reconciled by hand across systems and start existing natively in the same graph. The three pieces that make up Lakeflow Lakeflow Connect covers ingestion, with point-and-click connectors for SaaS applications (Salesforce, Workday, ServiceNow), databases (SQL Server), and messaging, plus Zerobus Ingest, a serverless direct-write API for anyone with events coming from their own applications who doesn't want to stand up a message bus in the middle. Spark Declarative Pipelines is the transformation layer: instead of hand-writing the orchestration of streaming tables and materialized views, you declare the target table and the business logic, and the engine infers the dependency graph on its own, builds the DAG, and takes care of operational details like backfill and versioning. The name has changed over time (anyone who has followed the platform for a while will recognize it as the direct evolution of what used to be Delta Live Tables), but the underlying mechanism, declaring the desired result instead of the procedural step-by-step, remains the central idea. Lakeflow Jobs orchestrates all of it as a unified DAG: SQL workloads, Python code, declarative pipelines, dashboards, and external systems live in the same job definition, with data-driven triggers (a table was updated, a file landed in a volume), control-flow tasks, and no-code backfill runs. Hands-on: a simple end-to-end pipeline To make it concrete, here's how the three pieces fit together in a basic ingestion and declarative transformation pipeline: import dlt from pyspark.sql.functions import col # Streaming table: incremental ingestion via Auto Loader @dlt.table( comment="Raw order ingestion via Auto Loader" ) def orders_bronze(): return ( spark.readStream.format("cloudFiles") .option("cloudFiles.format", "json") .load("/Volumes/sales/bronze/orders_raw") ) # Materialized view: validation and enrichment @dlt.table( comment="Validated orders, with a quality filter" ) @dlt.expect_or_drop("positive_amount", "total_amount > 0") def orders_silver(): return ( dlt.read_stream("orders_bronze") .withColumn("total_amount", col("total_amount").cast("decimal(10,2)")) .filter(col("customer_id").isNotNull()) ) This snippet doesn't run on its own, it becomes a task inside a Lakeflow Job that can also kick off an AI/BI dashboard as soon as the "orders_silver" table is updated, using a data-driven trigger instead of a fixed time-based schedule. That binding between ingestion, transformation, and the next step in the flow, all in the same place, is what eliminates much of the glue work that today lives in a separate Airflow or Azure Data Factory. Governance without manual reconciliation The point that usually goes unnoticed by anyone looking only at the product surface is what happens underneath with Unity Catalog. Because ingestion, transformation, and orchestration share the same catalog, data lineage is captured end to end automatically, from the source table in SQL Server all the way to the final dashboard, without relying on manual annotation or on a third-party lineage system trying to infer relationships from logs. My take: this unification of governance is, in my view, Lakeflow's strongest argument, stronger even than the convenience of having a ready-made Workday connector. I've seen more than one data engineering project fail an audit not because the pipeline was wrong, but because nobody could prove with confidence where a specific number had come from, and how many transformations sat in between. Having that guaranteed structurally, rather than as a manual documentation process that is always out of date, changes the kind of conversation you have with the compliance team. System Tables come in as the observability piece: instead of building your own monitoring dashboard by querying each separate tool's API, you can build alerts and health reports directly on top of native SQL tables, with retention and schema standardized by Databricks itself. Cost: where the promise requires your own verification Databricks publishes some impressive cost-reduction figures, including a case of up to 83% lower ETL cost and up to 90x performance improvement reported by a specific customer. Those numbers deserve the standard skepticism any vendor benchmark does: they usually reflect a specific migration, from a poorly optimized stack to a new, well-tuned configuration. In practice, I'd test the real savings by running the same representative workload from your own environment on old vs. new, with equivalent cluster policies and compute mode (Performance vs. Standard on serverless compute), before using the marketing number in a budget justification to leadership. The mechanism behind the savings is real and makes technical sense: serverless compute with automatic optimization eliminates cold starts between sequential tasks in the same job, and granular per-task resource control avoids over-provisioning a cluster for a lightweight stage of the pipeline. That's different from guaranteeing that any migration will hit an 83% reduction. What this doesn't solve Consolidating ingestion, transformation, and orchestration into one platform doesn't remove the need for well-thought-out data modeling, nor does it replace architecture decisions about partitioning and clustering large tables. It also doesn't, on its own, solve the problem for teams that already have a heavy investment in Airflow with custom plugins that are hard to port, migrating legacy orchestration is still manual rewrite work, even with the dbt task and third-party integrations easing part of the path. And Lakeflow Connect's point-and-click connectors cover a specific set of sources, legacy systems or proprietary APIs without a native connector still depend on custom ingestion via Auto Loader or your own API. Bottom line Lakeflow doesn't invent a new data engineering paradigm, it removes the friction of coordinating three different tools to do what was always conceptually a single thing: bring data in from outside, transform it with confidence, and deliver it in the right shape to the next consumer. In practice: for anyone on Azure Databricks who already feels the pain of maintaining Data Factory, a third-party orchestrator, and a separate lineage catalog, it's worth piloting on a real pipeline before promising Databricks' cost-reduction number to anyone. References Databricks Blog, "Modernize your Data Engineering Platform with Lakeflow on Azure Databricks": https://www.databricks.com/blog/modernize-your-data-engineering-platform-lakeflow-azure-databricks Databricks Docs, "Get started: Build an ETL pipeline": https://docs.databricks.com/aws/en/getting-started/data-pipeline-get-started Microsoft Learn, "Tutorial: Build an ETL pipeline with Lakeflow pipelines - Azure Databricks": https://learn.microsoft.com/en-us/azure/databricks/getting-started/data-pipeline-get-started Microsoft Learn, "System tables": https://learn.microsoft.com/en-us/azure/databricks/admin/system-tables/52Views0likes0CommentsBurst IBM Spectrum Symphony workloads into Azure with the new Azure Compute Fleet provider plugin
A new resource connector links Symphony Host Factory to Azure Compute Fleet, so established Symphony environments can elastically extend into Azure. Why this matters for large-scale compute Enterprises running large scale AI, analytic, and high-performance computing (HPC) environments need more than infrastructure. They need to dynamically match thousands of jobs with available capacity while balancing performance, resilience, availability, and cost. Microsoft Azure and IBM are announcing the IBM® Spectrum Symphony Provider Plugin for Microsoft Azure Compute Fleet. The integration enables IBM Spectrum Symphony and Azure Compute Fleet to dynamically access and manage Azure compute capacity, providing an optimized execution environment and allowing enterprises to extend and elastically burst established Symphony environments to Azure. IBM Symphony is a proven workload orchestration platform used across financial services, manufacturing, life sciences, energy, and other compute-intensive industries. How the integration works At the center of the integration is an Azure resource connector for IBM Host Factory. The connector translates Symphony resource requirements into Azure Compute Fleet requests, creating a direct bridge between the Symphony orchestration layer and Azure infrastructure. The architecture maintains a clear separation of responsibilities: Symphony manages the workload while Azure Compute Fleet manages access to compute capacity. The request path is straightforward: When workload demand increases, Symphony determines the additional resources required through Host Factory. The Azure resource connector translates those requirements into a Compute Fleet request, and Azure provisions eligible infrastructure. The resulting Azure virtual machines join the Symphony environment as execution hosts and begin processing jobs. As workload demand decreases, the resources can be returned. What the plugin delivers The Microsoft and IBM collaboration brings together Symphony's enterprise workload orchestration capabilities with the scale and flexibility of Azure Compute Fleet: Dynamically provision and scale as Symphony workload demand increases or decreases, keeping resource use matched to actual need. Maintain workload performance through policy-driven elasticity governed by existing workload priorities and service- Take advantage of Azure Virtual Machines' wide array of SKU choices by flexibly deploying based on workload requirements and available capacity. Seamlessly leverage Spot instances, so that workloads can run at optimal cost. Support hybrid workload orchestration that extends established Symphony environments to Azure, realizing the cost and availability benefits of hybrid multicloud deployment. Capabilities available to Symphony customers on Azure Symphony on Azure enables customers to: Access a broader pool of Azure capacity by provisioning across multiple eligible VM families instead of relying on a single VM SKU. Select infrastructure using workload attributes such as vCPU, memory, storage, and other requirements. Balance availability and cost through a flexible combination of Azure Spot Virtual Machines and pay-as-you-go VMs. Increase workload throughput and infrastructure utilization by matching queued jobs with suitable Azure capacity. Expand compute capacity during periods of peak demand while provisioning steady-state infrastructure for typical usage. These capabilities are particularly valuable for financial risk calculations, scientific computing, engineering analysis, and other large-scale batch workloads. Different jobs can require different combinations of CPU, memory, performance, availability, and price, and Azure Compute Fleet can provision across eligible VM types and purchasing models, expanding the pool of infrastructure available to execute those workloads for faster, more cost-efficient deployments. Bringing established environments forward Many enterprises have invested years developing sophisticated Symphony environments with application-specific scheduling policies, priorities, queues, service-level objectives, and operational processes. The plugin preserves that investment: scheduling policy and operational process stay in Symphony, while Azure supplies elastic capacity underneath. Microsoft and IBM are working with joint enterprise customers across financial services and other compute-intensive industries to expand adoption and advance the integration. "Symphony's native enablement on Azure gives our joint customers a best-in-class platform for orchestrating HPC workloads with the performance, resilience, and cost efficiency they require," said Sripriya Srinivasan, General Manager, IBM Software Products, IBM. "This is a transformational step forward for hybrid computing." Partner opportunity For partners with HPC, risk analytics, or scientific computing practices, this opens repeatable engagement motions: assess existing Symphony estates, design the Compute Fleet capacity strategy across Spot and pay-as-you-go, package burst-to-Azure as a managed offer, and grow consumption as workloads scale. Get started Read the full technical walkthrough on how to burst IBM Spectrum Symphony into Microsoft Azure with the new Azure Compute Fleet cloud provider: Microsoft and IBM bring enterprise-scale workload orchestration to Azure with IBM Spectrum Symphony | Microsoft Community Hub To download the Azure Compute Fleet plugin for IBM® Spectrum Symphony, visit the Azure Marketplace.124Views0likes0CommentsAzure Databricks Cost and Workload Assessment: a Local Web UI
Introduction To review an Azure Databricks environment, you need to know what it costs, how its compute is used, and which jobs need attention. That information is spread across Azure billing and Databricks. The Assessment & Optimization Workbench is a local web application that brings these details together. Select your workspaces and dates, check access, and run an assessment. You can then explore the results in your browser and download a report to discuss with your team. The walkthrough below shows each step. 1. What you can assess The assessment covers costs, workload behavior, and selected configuration checks: Area What you can inspect Cost Actual and Amortized costs, top drivers, attribution gaps, and commitment scenarios with eligible hourly inputs. Efficiency CPU/memory samples, idle observations, sizing candidates, query duration, and queueing. Job health Failures, retries, failure notifications, and network observations. Inventory and posture Workspace assets, IP access list settings, and Unity Catalog metastore assignment. Evidence Collection coverage, snapshots, raw-data imports, and offline re-analysis. Handoff Prioritized findings, optional review decisions, formatted reports, and Excel exports. If some data could not be collected, the UI shows what is missing. Recommendations are suggestions for you to review and test; the assessment does not apply them automatically. Actual savings must be measured after you make the recommended changes. 2. The five-step workflow The UI guides you from choosing what to assess to downloading a report. First, select your workspaces and check access. Then start the assessment, explore the results, and export what you need. The walkthrough below shows each step. Configure → Validate → Run analysis → Visualize results → Review & export Configure This Configure step defines what the assessment will cover: which workspaces to review, the date range, and the data to collect. These choices prepare the assessment; they do not start collection. The GIF shows the scope selection, warehouse choices, and date settings before moving to validation. Select the scope: choose the subscriptions, resource groups, and Databricks workspaces to include. Set the dates and cost basis: choose the period to assess. ActualCost shows charges as recorded; AmortizedCost spreads eligible commitment costs over time. Choose the collection profile: use Standard for core assessment data, Extended to include workspace asset metadata, or Custom to select optional data and analysis modules. Select SQL Warehouses: choose a warehouse for each workspace that needs system-table queries. Selection does not start compute. Queries can incur charges, so warehouse use requires approval in Validate. Next: Select Validate configuration. Validate Validation checks your setup and access to the selected Azure and Databricks data. It helps you identify missing permissions or approvals before starting the assessment. Select Run validation, approving warehouse auto-start if required. The UI shows which checks are running, which have passed, and what needs attention. The GIF shows an approval issue being resolved, followed by the access checks. Fix any blockers that prevent the assessment from running, and read warnings about data that may be limited or unavailable. If permissions are missing, the permission panel provides commands for an authorized administrator to run. Validation itself does not grant access. Next: Once validation allows you to proceed, select Continue to run below the permission panel. You will review the setup and start the assessment separately. Run analysis This step collects data from your selected Azure and Databricks workspaces, analyzes it, and creates the assessment report. You can follow the progress and see which sources returned data. Review the selected workspaces and date range, then select Start read-only assessment. The screen shows: Progress: the current stage, from initial checks and collection through analysis and report generation, plus elapsed time. Collection sources: cards grouped into Azure data and Databricks workspaces. Each card shows its status and the number of items collected, such as cost records, inventory, or billing data. Console: detailed messages to help explain delays or collection problems. When the run finishes, green means collection and analysis succeeded, not that every workload is healthy. Red means the run failed or some data needs attention. A skipped source was not collected, often because an optional check was not selected. The report and collected data are saved locally, including any reported gaps. Next: Select Visualize results to explore the saved assessment. Visualize results This step turns the collected data into charts, tables, and findings you can explore. Start with the overall cost and collection summary, then look at individual workspaces, compute resources, and jobs to understand what needs attention. You are viewing saved results, so switching tabs does not collect data again. The GIF follows the results from cost and workload details through findings, data quality, and a proposed roadmap. The tabs are grouped into Overview, Technical, and Decisions. Overview: understand the cost Executive summary: see total cost, spending mapped to workspaces, optimization candidates, and collection status. Check whether the detailed costs match the billing total and how much supporting data is available. Cost analysis: use Cost evidence to compare Actual and Amortized trends. Break down spending by Service, Meter category, SKU, Resource group, Workspace, or Owner tag, then inspect the largest cost drivers and unmapped spend. Commitment opportunities lets you enter a proposed number of committed nodes and calculate a scenario when the required hourly usage and pricing data is available. It does not purchase a commitment. Technical: inspect resources and workloads Compute and SQL has five views: Inventory: cluster and warehouse settings, including node types, worker counts, Photon, and automatic shutdown, plus job run and failure summaries. Utilization: CPU, memory, idle observations, and sample counts. Select a resource and a driver or worker instance to view its CPU and memory chart. Sizing: current node and worker configuration, with candidates to benchmark before resizing. Job health: runs, failures, notification settings, tasks, and retry policies. Network: data sent and received by nodes, plus CPU-wait measurements. These are traffic observations, not billed network charges. Queries: switch between Individual queries, Warehouse summaries, and User summaries. Review durations, queue times, and failures; search and sort to find queries that need investigation. Posture: inspect IP access list settings and Unity Catalog metastore assignment. Each check shows its observed value and outcome; this is not a full security audit. Assets: browse collected repositories, notebook metadata, MLflow experiments, serving endpoints, SQL alerts, Genie spaces, and Unity Catalog volumes. These appear only when the relevant optional data was collected. Decisions: check findings and plan the next steps Findings: browse recommendations by category, status, and confidence. Select a row to read what was observed, the recommended next step, supporting records, and any limitations. Evidence quality: see coverage by workspace, missing data, failed or skipped sources, and records excluded from scope. Open Inspect effective rules / create another analysis to review or change analysis settings, then Create child analysis to rerun them on saved data without changing the original snapshot. Roadmap: review proposed work across Days 0-30, 31-60, and 61-90, including owners and dependencies. The measurement plan explains the baseline for checking savings after a change; the roadmap is not an approved implementation schedule. Use Filters to narrow findings by subscription, resource group, workspace, workload, category, confidence, or status. Scope selections also apply to supported technical tables, which have their own search, sorting, and paging controls. Select a resource name to inspect its saved details. If data is unavailable, check Evidence quality before drawing a conclusion. A missing measurement does not mean that a resource was unused. Next: Select Review & export to read the report or download the results. Review & export This final step lets you read the assessment report and download files to discuss with your team. You can also record decisions on individual findings. Review is optional, so you do not need to approve every finding before exporting. The GIF shows a finding being reviewed, the formatted report and supporting files being opened, and an Excel workbook being generated and downloaded. Read the report: select Preview report to move directly to the formatted report in your browser. Use its contents links to jump to sections, and open supporting evidence links to inspect the saved files. Download the report: select Download report to save the original Markdown version. Downloading a file completes this workflow step, but does not approve findings or apply changes. Record a decision: expand Record a decision, choose a finding and decision, then enter the reviewer and an optional note. Select Save review decisions to retain the changes. A decision other than pending requires a reviewer. Inspect supporting files: use Run artifacts to preview supported files or download individual outputs. Reports appear as formatted pages; CSV and JSON previews show the file contents as text. Create an Excel workbook: choose the modules under Excel workbook, then select Generate workbook artifact and download the generated file. It includes summary, rules, quality, findings, saved review decisions, and the selected module sheets. It covers the full saved assessment, not just the rows currently filtered in the UI. Save or discard pending review edits before generating a workbook. Exports include saved decisions only. Check the files before sharing them because supporting evidence can contain sensitive environment details. Optional dashboard publication is a separate cloud action for publishing coverage counts, not the full assessment. It requires a destination workspace and warehouse, a preview of the publication plan, and explicit approval of the write and possible warehouse charges. 3. Reopen saved assessments A snapshot is a locally saved assessment, including the selected scope, collected data, findings, reports, and review decisions. It lets you return to earlier results, continue a review, or download files later without querying Azure or Databricks again. Reopen an assessment: choose a run from Saved snapshots. Entries show the run date, customer, status, and run ID, with the newest first. You can then explore its results or continue in Review & export. Try different analysis settings: open Evidence quality, expand Inspect effective rules / create another analysis, adjust the settings, and select Create child analysis. This creates a separate assessment using the same saved data. The original remains unchanged, and the new findings need their own review. Remove old assessments: use Manage snapshots to delete individual snapshots or clear the saved history. Deletion requires confirmation and permanently removes the selected runs, including their evidence, reports, and review decisions. Back up important runs first. Snapshots show what was collected at the time, not the current environment. To get newer data or collect a missing source, start a new assessment. 4. Architecture and collection flow The workbench runs on your machine, with a browser interface connected to a local Python server at http://127.0.0.1:8765. Azure and Databricks supply the data; collection scripts, analysis, and saved results stay local. No Azure-hosted application is needed. The diagram follows an assessment through seven stages: Configure: the React and TypeScript browser UI captures your selected workspaces, dates, and collection options. Coordinate: the Python API receives requests from the browser, launches assessment processes, and tracks progress. Check access: validation checks the selected scope, permissions, and required approvals. These checks use live access, but full collection starts only when you select Start read-only assessment. Collect: PowerShell scripts read Azure resource inventory and billing data, plus Databricks APIs and system tables. Requests to these external sources use HTTPS. Save evidence: each run stores the returned data, configuration, source outcomes, and logs in a local folder. Analyze: Python combines the saved data, checks costs against billing totals, identifies data gaps, and applies rules to produce findings and reports. Review and export: the API returns saved results to the browser, where you explore findings, record decisions, and download reports or data files. Keep in mind: the assessment does not apply recommendations. SQL Warehouse queries can incur charges, and exported files should be checked for sensitive details before sharing. 5. Get started Try it Prerequisites: PowerShell 7+, Python 3, Node.js 22.12+/npm for the build, and Azure CLI signed in. Collection needs Azure Reader/Cost Management Reader and the required Databricks source access. From the repository root: Set-Location .\ui npm ci npm run build Set-Location .. .\ui\Start-AssessmentUi.ps1 Open http://127.0.0.1:8765 and keep the launcher running. If already built, only the final command is needed. Learn more User guide Required permissions Test record and limitations Contribute Report issues in the repository with reproduction steps and redacted source statuses. Never include credentials or raw customer evidence. Publishing check: redact environment identifiers in GIFs and upload media when posting outside the repository. Displayed timings and amounts are not benchmarks or savings claims.135Views0likes2Comments"Authorization failed" error for Logic app writing a comment to Sentinel Incident
I have created a managed identity named id-sentinel-playbook that is used in 2 logic apps. Both the logic apps retrieve information from different external apis and writes the results as comments into the Sentinel incident. The managed identity id-sentinel-playbook has been assigned 2 roles - Microsoft Sentinel Responder and Microsoft Sentinel Automation Contributor role (See screenshot). However when one of the logic apps transacts with Sentinel such as checking the watchlist or writing comment into a Sentinel incident, there is the 403 forbidden error (See screenshot). It works fine when I use my Azure account as connection for the logic app. The other logic app also works fine when the same managed identity id-sentinel-playbook is used as connection to Sentinel. I have compared the identity of both the logic apps and they are the same. I have also searched online for existing answers and all point to the managed identity having insufficient roles, however id-sentinel-playbook already has the Microsoft Sentinel Responder role and strangely the other logic app that writes comments into the Sentinel incident as well, works. Here is the screenshot of the logic app having the user managed identity. The other logic app has the same. Please help. I spent 2 days investigating this and have no more ideas on how to further investigate this😓.592Views0likes2CommentsHow to update the proxyAddresses of a Cloud-only Entra ID user
I currently have a client with an Entra ID user (not migrated from on-premises) that is cloud-based, but has proxyAddresses values assigned. Now, I want to update the proxyAddresses through the Graph Explorer and have used this link as a guide: https://learn.microsoft.com/en-us/answers/questions/2280046/entra-connect-sync-blocking-user-creation-due-to-h. Now this guide is suggesting you can use the BETA model and this URL format... https://graph.microsoft.com/beta/users/%USERGUID% It states you can use that URL to do both 'GET' and 'PATCH' queries - the PATCH query being the one that will change the settings. You have to put forth a body for the proxyAddresses property in the PATCH query, which represents all of the addresses you want the user to utilise as proxy addresses. Now the GET query works... The PATCH query does not... Screenshot provided: Now, regarding the error message, I have applied ALL possible permissions in the 'Modify Permissions' tab. It is still erroring, Now I cannot use Exchange Online PowerShell, as the user does not have a mailbox! Aside from potentially using a license for Exchange Online or provisioning a mailbox for the user, and making the necessary changes, would the only other option be to delete/recreate the user?Solved1.7KViews0likes4Comments