DataBricks Archives -

Category Archives: DataBricks

Building a 3-Layer Pipeline to Cleanse Business Central Data for Forecasting

Posted On September 28, 2026 by Posted in

Summary This article explains how a data pipeline can be used to prepare Business Central data for demand forecasting. Raw data may contain missing values, inconsistent fields, and information that is not immediately suitable for analysis. The pipeline uses a three layer Medallion architecture: Bronze for storing the raw data, Silver for cleaning and preparing it, and Gold for creating the final data needed for analysis and forecasting. A key principle throughout the process is “NULL is not 0” . A missing value does not always mean that the value is zero. The pipeline therefore uses context aware handling and missing data flags so that the original meaning of the data is not lost. The article explains the complete flow from data ingestion through cleansing, NULL handling, transformation, and aggregation, along with the code and examples used to implement each step. Table of Contents 1. Introduction 2. The Business Problem The Solution 3.1 Medallion Architecture Overview 3.2 Ingestion API from Business Central 3.3 The NULL Philosophy 3.4 Silver: Intelligent Imputation 3.5 Flagging the Unknown 3.6 Gold: Aggregating Without Lying 4. Limitations & Considerations 5. Business Impact 6. FAQs 7. Conclusion 1. Introduction In modern supply chain and demand forecasting, data quality is the foundation of every decision. Business Central can serve as a source for sales, inventory, purchases, and transfer data. However, raw CDC (Change Data Capture) payloads extracted through the OData API may require additional cleansing before they are forecast ready. Fields were missing, column names were inconsistent, and most critically NULLs were everywhere . While Business Central stores transactional data reliably, the OData API exports often include null values for optional fields, unpopulated attributes, or fields that simply don’t apply to certain record types. Treating these NULLs as zeros would have been a catastrophic mistake in demand forecasting. A NULL in a sales quantity column doesn’t mean “zero units sold” it means “we don’t know”. Erasing that distinction would hide the difference between genuinely zero demand and missing data, leading to under forecasting, misplaced safety stock, and costly stockouts. To solve this, I built a 3 layer Medallion pipeline (Bronze → Silver → Gold) in Azure Databricks. I also designed a custom ingestion API using Databricks’ OData V4 connector wizard to pull 12 ERP entities directly into Azure Blob. The Silver layer introduced a “NULL is not 0” philosophy, with context aware imputation and boolean flags to preserve data quality signals. The Gold layer aggregated daily demand per item location, leaving NULLs as NULL to correctly represent “no data” days. This article walks through the architecture, the imputation logic, and the practical lessons learned from transforming 12 messy datasets into a clean, forecast ready mart. 2. The Business Problem The organization relied on manual Excel spreadsheets to prepare demand forecasts across 22 high impact items and 3 warehouse hubs a total of 66 item location combinations . This manual process was error prone and unscalable. Several challenges emerged: Heavy reliance on spreadsheets: Demand planning was executed via disconnected, static Excel files. Frequent formula errors, broken references, and accidental overwrites led to incorrect estimates. Lack of scalability: Manually calculating forecasts across 66 combinations was time consuming. Excel could not dynamically model complex interactions like monthly seasonality and day of week demand patterns simultaneously. Reactive procurement and stockouts: Teams spent more time cleaning data than analyzing risk. Static buffer stocks failed to account for changing supplier lead times, resulting in emergency purchase orders or stockouts during peak seasons. NULLs were treated incorrectly: In many cases, missing values were simply replaced with 0, masking the true state of the data and causing the forecast model to underestimate demand. The Objective: Build a reliable, automated pipeline that ingests raw CDC data from Business Central, cleanses it with context aware logic, and produces a daily demand mart that distinguishes between “zero sales” and “unknown sales” . 3. The Solution The solution was designed as a Medallion architecture with three distinct layers, each serving a specific purpose in the data quality journey. 3.1 Architecture + Diagram 3.1 Medallion Architecture Overview The pipeline processes 12 raw JSON datasets from Business Central through Bronze, Silver, and Gold layers in Azure Databricks. SVG Diagram 1: Medallion Architecture Medallion Architecture Business Central Data Pipeline Business Central OData V4 / REST API 12 CDC datasets Databricks OData V4 Connector Auth · Pagination · Parsing Azure Blob Bronze (Raw JSON) Immutable landing 🥈 Silver Layer Cleaning · Typing · Flagging • Remove @odata.etag • Standardize columns · Cast types 🥇 Gold Layer Aggregation · Business Logic • Daily demand per item location • NULLs preserved as “no data” 📊 Forecasting & Analytics Power BI · ML Models · Business Decision Support Read only pipeline cleaned data remains in the analytical pipeline Medallion architecture: Business Central OData → Bronze → Silver (cleaning) → Gold (aggregation) → Analytics 3.2 Ingestion API from Business Central The first step was to build a reliable ingestion layer. I used the OData V4 connector wizard in Databricks to set up an authenticated connection directly to Business Central’s REST API. This connector handled pagination, authentication (Azure AD OAuth 2.0 with Key Vault), and parsing of the OData response format (data wrapped inside a “value” array). I configured it to pull all 12 datasets that feed into the forecasting pipeline. Security: Zero hard coded credentials all secrets stored in Azure Key Vault. Endpoints: Pulled items , sales_lines , purchase_headers , item_ledger_entries , and 8 other entities. Landing: Raw JSON payloads were written to Azure Blob (Bronze) as immutable snapshots, preserving the original CDC state. 3.3 The NULL Philosophy Before writing a single line of PySpark, I defined a clear rule: NULL is not 0 . This principle guided every transformation in the Silver layer. In demand forecasting, a NULL in a sales quantity column means “we don’t know”. It could be due to a missing transaction, a CDC gap, or a field that doesn’t apply to that record type. Treating it as zero would artificially deflate demand, … Continue reading Building a 3-Layer Pipeline to Cleanse Business Central Data for Forecasting →

Share Story :

How CloudFronts Built a Project Risk Assessment Engine with Databricks Genie

The Warning Signs Were Always There – Project Risk Engine Summary On most delivery projects, the warning signs appear well before the escalation. They sit in email threads, support case notes, overdue invoice reminders, and resource utilization reports – scattered across systems that nobody has the time to cross-reference every week. By the time a project gets formally marked at risk, it is usually already a difficult conversation. This post shares what we built at CloudFronts to solve that problem from the inside. The Project Risk Engine uses a multi-model AI approach running on Azure Databricks to read unstructured project communication, score project health, and flag projects in RED – all surfaced to project managers inside Microsoft Teams through Databricks Genie, in plain English. We built it for our own PMO first. This article walks through why we built it, how it works, and what it changed. Table of Contents 01 Where This Started 02 Where Project Health Actually Lives 03 What the Risk Engine Does 04 How It Works – Four Layers 05 How a Project Gets Flagged RED 06 Genie in Teams 07 What Changed for the Team 08 Who It Helps 09 How to Implement This 10 Challenges We Faced 11 FAQs 12 Conclusion Where This Started This did not start as a product idea. It started as a recurring frustration inside our own team – the group at CloudFronts responsible for delivery and billing excellence. Three things kept happening. Early warning signs were buried in long email chains, so risks only surfaced after they had grown. Project managers had no quick way to assess project health, which meant the team reviewed each project by hand every week just to work out where things stood. And because we run many projects simultaneously, it was easy for a real problem to slip through unnoticed. Here is a real example. One of our clients had a support agreement coming up for renewal. It is the kind of thing everyone assumes someone else is tracking – but it was not flagged until almost the last day. If that client had decided not to renew, leadership would have had no early warning at all. The information was there. It just was not in front of anyone in time. In another case, a project had two phases and the client had not yet confirmed the scope for the second. That was a clear risk to the final payment. Most of us were aware of it in some abstract way – but the project was not formally flagged until the invoice was already due. A few weeks earlier would have made a real difference. “The warning signs were already there. They were just scattered, and easy to miss until it was too late.” Where Project Health Actually Lives A project’s health is not in one place. It is spread across four areas – and any one of them, left unmonitored, can become a risk that affects cash flow, client relationships, or delivery quality. Booking Contract renewals, new work coming in, and MSA status. A contract quietly approaching expiry is a risk most teams only notice when it is almost too late. Billing & Delivery Project status, open risks, support tickets, internal notes, and the client emails tied to them. This is where the unstructured signals sit. Collection What has been invoiced, what is outstanding, and what is running overdue. Delayed invoicing and unpaid milestones are early indicators of a project in trouble. Resource & Utilization Who is working on what, how busy they are, and where allocation gaps are forming. Low utilization on a contracted engagement is a billing risk waiting to happen. Any one of these on its own does not tell you much. You only see the full picture when you put all four together – and that is exactly what the engine does. What the Project Risk Engine Does 1Unifies the data – Delivery, support, billing, and resourcing data from Dynamics 365 CRM and Outlook is pulled into one governed dataset on Azure Databricks, structured through a Medallion Architecture (Bronze → Silver → Gold). 2Reads the unstructured signals – AI reviews client emails, internal notes, and case history to surface hidden threats: delays, blockers, escalation patterns, and unanswered client requests that never make it into a status field. 3Scores and flags – Every project receives a health score from zero to ten, along with a risk summary, identified threats, and next recommended actions. Projects meeting RED criteria are automatically flagged. 4Makes it conversational – Project managers can ask questions about any project in plain English, directly inside Microsoft Teams through Databricks Genie. No dashboards to open, no reports to pull. How It Works – A Multi-Model, Four-Layer Approach The risk analysis is not a single AI call. It is a deliberate, layered process – lightweight checks first, heavy AI only where it is needed. This keeps it fast and cost-efficient while ensuring accuracy on the things that matter. Layer1 Rule-Based Pre-Filter No AI involved. Simple rules filter out inactive email threads not part of any recent communication, eliminating noise before any model is engaged. Layer2 Lightweight AI Triage A smaller model sorts remaining threads into four categories: active, completed, informational, or waiting. Only active and waiting threads move forward. Layer3 Deep Risk Analysis A larger model – Claude Sonnet 4.5 – runs deep analysis on threads that survived the first two layers. It is used here because this stage requires stronger reasoning: detecting underlying threats, assessing likelihood, and estimating impact. Layer4 Executive Summary Everything is rolled up into a project health score (0–10), top threats, and recommended next actions – a clear, evidence-backed picture of where each project stands. Solution Architecture Here is how the full architecture fits together – from data sources through to the project manager asking a question in Teams. Data flows from Dynamics 365 CRM and Outlook → Azure Logic Apps → ADLS Gen2 Medallion layers → Azure Databricks → Genie AI → Microsoft Teams … Continue reading How CloudFronts Built a Project Risk Assessment Engine with Databricks Genie →

Share Story :

From Project Reporting to Project Intelligence: How AI is Transforming Project Management

Summary We built a Databricks Genie agent for our own PMO at CloudFronts, running on Dynamics 365 data held in a Databricks lakehouse. Project managers ask a question in plain English and get an answer back across resource utilization, time tracking, billing and milestones, tickets and cases, and project status. This blog covers what the agent does, what email sentiment analysis shows that the numbers do not, and how it works inside Microsoft Teams. Table of Contents Introduction The Challenge The Solution See It in Action Business Impact Frequently Asked Questions Conclusion Introduction This started inside our own PMO — the Billing and Delivery Excellence function at CloudFronts. We learned about a project risk when someone escalated it. The warning signs came earlier than that, in email threads and internal notes, but reading every thread across every project each week was not work anyone could take on. The rest of the picture was split across systems. Billing held the invoice that had passed its due date, delivery held the milestone that had moved, support held the ticket that had been open for weeks. No one screen put those next to each other, so the PMO opened each project every week and compiled the status by hand. So we built the agent for ourselves first: a Databricks Genie agent running on Dynamics 365 data held in a Databricks lakehouse, which project managers query in plain English. The Challenge Dynamics 365 Project Operations holds everything a project manager needs — resource assignments, logged hours, billing milestones, project budgets, and delivery timelines. The data is there. The challenge is that getting specific answers from it still requires navigating multiple modules, running reports manually, and in many cases, exporting to spreadsheets to piece things together. This created a set of questions that were surprisingly hard to answer: Identifying which resources are overutilized or sitting idle requires pulling allocation data and comparing it manually against actual hours logged Understanding whether a project is at risk means cross-referencing milestone progress, budget consumption, and team capacity — a process that can take hours Billing questions — what has been invoiced, what is pending, what is approaching a milestone — require moving between finance and project views that are not always aligned Status updates for leadership need to be manually compiled, often pulling from data that was accurate yesterday but may have shifted today The result is that project managers operate on a lag — making decisions based on reports that reflect the past, not the present, and spending time producing those reports instead of acting on them. The Solution — A Genie Agent Built on Databricks and D365 Project Operations We built a Genie agent on Azure Databricks, connected to Dynamics 365 Project Operations. Project managers can now ask questions in plain English and get answers drawn directly from their project data — without building a single report. The agent is designed around the areas that matter most to project managers on a daily basis: a. Resource UtilizationThe agent can answer questions about who is overallocated, which resources have capacity available, and how utilization is trending across the team or a specific project. What previously required pulling allocation reports and comparing them against timesheets can now be answered in a single question. b. Time TrackingProject managers can ask which team members have not logged hours for the week, where hours are being spent versus what was planned, and whether a specific project is tracking within its estimated effort. The agent surfaces this from logged timesheet data in D365. c. Billing and MilestonesThe agent connects billing milestone data with project progress, allowing project managers to ask what is due for invoicing, which milestones are approaching, and whether any billing triggers are at risk of being delayed. This brings finance and delivery into the same conversation. d. Tickets and CasesThe agent surfaces open tickets and cases linked to a project — how many are open, which are overdue, how they are distributed across team members, and whether any are blocking delivery. Project managers can ask for a snapshot of issue health across one or multiple projects without navigating case queues manually. e. Email Sentiment AnalysisOne of the more telling signals of how a project is going is often hiding in the inbox. The agent analyses email communication patterns and sentiment across project stakeholders — flagging when tone is shifting, when a client’s responses are becoming shorter or more urgent, or when concerns are being raised repeatedly. This gives project managers an early, qualitative read on relationship health before it shows up in a formal escalation. f. Project StatusInstead of assembling a status report, a project manager can ask for a summary of where a project stands — budget consumed, milestones completed, risks flagged, and remaining timeline. The agent compiles this from D365 data and presents it in plain language, ready to share or act on. The conversation does not stop at one question. A project manager can ask a follow-up — drill into a specific resource, filter by project phase, or compare two projects side by side — and the agent follows the thread, refining its response at each step. Available directly in Microsoft TeamsThe Genie agent is also available as a Databricks App inside Microsoft Teams — meaning project managers do not need to switch tools to get answers. They can ask questions about their projects, resources, and billing directly from the Teams interface they already work in every day. See It in Action Weekly Work Summary — Time Tracking in ActionA project manager asks Genie for a summary of work completed last week. The agent returns a full breakdown — total hours logged, billable vs non-billable split, project-wise distribution, and key observations — in seconds. Case Detail View — Tickets and Cases in ActionA project manager asks for details on a specific case. The agent surfaces the full case record — status, owner, priority, activity timeline, and a follow-up alert — without the manager needing to … Continue reading From Project Reporting to Project Intelligence: How AI is Transforming Project Management →

Share Story :

From ERP Data to Process Mining Insights: Building an Automated Pipeline for Real-Time Process Visibility

Summary Clean ERP data sitting in a data lake doesn’t answer the question every operations leader eventually asks: where exactly is our process breaking down? We built an automated pipeline that connects a client-facing web portal, Azure Table Storage, and Azure Databricks to a leading process mining platform, turning validated ERP data into a living view of how work actually flows. The pipeline is fully status-driven: every record is tracked from submission through processing to completion, with no manual exports or spreadsheet hand-offs. Purchase order data is modeled through a medallion architecture and delivered to the process mining platform, where AI-driven analysis automatically surfaces bottlenecks and deviations from the expected process. Business impact: process owners moved from static, after-the-fact reporting to a near real-time, evidence-based view of process performance. Table of Contents 01  About the Customer 05  The Six-Step Pipeline 02  The Challenge 06  Architecture Overview 03  The Solution 07  Business Impact 04  AI-Driven Process Mining 08  FAQs About the Customer Customer Spotlight A Leading Digital Transformation Partner — Europe Our customer is a leading enterprise headquartered in Europe, operating across diverse manufacturing and supply chain divisions. Having already standardized their ERP data through a medallion architecture on Databricks, leadership wanted to go a step further: not only manage ERP data at scale, but also connect it seamlessly into process mining tools to uncover how core processes truly perform in practice. The focus was on gaining operational clarity into workflows such as purchase order management, invoice handling, and procurement cycles. The Challenge Standardized, clean data answers “what happened.” It rarely answers “why is this taking so long” or “where exactly is this process breaking down.” The business kept running into the same limitations: 1Why do purchase orders take longer to close in some regions than others? 2Which approval step is quietly adding the most delay to the process? 3How do we get validated ERP data into a process analysis tool without manual exports every time? 4How do we know, at any point in time, what has been processed, what’s pending, and what failed? 5Can this insight be generated automatically, instead of requiring a manual investigation every quarter? The Solution We extended the existing Databricks-based data platform with an automated, status-driven delivery layer connecting a client web portal, Azure Table Storage, Azure Databricks, and a leading process mining platform, orchestrated end-to-end with minimal manual intervention. Status-Driven Orchestration Every record carries a live status, from initial submission through sync completion, tracked in Azure Table Storage. Automated Bulk Processing Azure Logic Apps trigger the pipeline through APIs, so batches of records are processed without manual intervention. Reusable Databricks Framework The same medallion pipeline used for data standardization models Purchase Order data for process mining. AI-Driven Process Analysis The process mining platform’s AI reconstructs the real, as-executed process and highlights bottlenecks automatically. The Six-Step Pipeline Here’s how a single record moves from submission to a fully synced, process-mining-ready state: ⚙ Client Web PortalEnd-to-end data pipeline · Azure + Databricks 6 steps 🌐 1) Website Input The user submits data via the client web portal, a form or API request initiates the pipeline. ↓ 🗃 2) Azure Table Sync Incoming data is written and synced into Azure Table Storage. ↓ 📁 3) Status Filter Records from Azure Table are filtered where status matches: ✓ Perfect🕑 Queue ↓ ⚡ 4) Databricks Pipeline The framework is executed through the Databricks pipeline, processing all filtered records in batch. ↓ 🔄 5) Azure Table Update Once the Databricks sync completes, status is updated in Azure Table: Queue→✓ Synced ↓ 📊 6) UI Reflection Synced data is reflected back to the client web portal UI for the end user. Architecture Overview Once records reach the “Synced” state, the same medallion architecture used for data standardization models Purchase Order Details and Purchase Order Lines and delivers them into the process mining platform: ERPExtracts Row-header files → Bronze Raw landing → Silver Cleansed & standardized → Gold Business-ready models → DeltaLake Parquet delivery → ProcessMining AI-driven analysis Because the framework is configuration-driven, the same architecture can extend to additional ERP data lake sources, SFTP feeds, or other cloud storage without a redesign. AI-Driven Process Mining Analysis With Purchase Order Details and Purchase Order Lines modeled and delivered on a reliable, automated cadence, the process mining platform’s AI reconstructs the real, as-executed purchase order process directly from the underlying event data. Instead of relying on assumptions about how the process should work, process owners see how it actually works: where orders stall, which approval paths deviate from the intended flow, and where cycle time is quietly being lost. “A purchase order may look fine on paper, but the process data tells you exactly where it got stuck, and that gap surfaces automatically.” Business Impact Before After Manual exports required to analyze process performance Fully automated, status-driven pipeline from intake to process mining No visibility into where a record stood in processing Live status tracking from submission through sync completion Process bottlenecks discovered through manual investigation AI-driven analysis surfaces deviations and delays automatically Static, after-the-fact process reporting Near real-time, evidence-based process visibility One-off integration effort per process area Reusable framework, extendable to other business processes Frequently Asked Questions Does this require a specific process mining platform? No. The pipeline delivers modeled, business-ready data through Delta Lake and Parquet, which can be connected to most modern process mining platforms. How often is data refreshed in the process mining platform? The pipeline is designed for batch processing on a defined schedule, and can be tuned toward near real-time delivery depending on business needs and source system constraints. Can this be extended beyond Purchase Order data? Yes. Because the framework is configuration-driven, the same approach can extend to other process areas such as order-to-cash or procure-to-pay. What happens if a record fails validation? Records that don’t meet the status criteria simply remain in a pending state and are not passed downstream, so failures are visible and traceable rather than silently dropped. Conclusion Clean data is the foundation, but process visibility is where the business … Continue reading From ERP Data to Process Mining Insights: Building an Automated Pipeline for Real-Time Process Visibility →

Share Story :

How a Self-Service Data Portal Solved Multi-Language and Domain Value Chaos in ERP Data

Summary Enterprises running large, multi-country ERP systems often extract data that is technically complete but practically unusable, split across duplicate language columns and encoded with undocumented numeric values. We built a self-service data platform on Azure so that business users, not just data engineers, could define, validate, and process ERP extracts without writing a single line of code. The solution resolves two of the most common ERP data problems: a single field like “Item Description” spread across nine language-specific columns, and reference fields like “Order Status” stored only as numeric codes. A custom web portal puts business users in control of table specifications, validation rules, and processing status, while Azure Databricks and Delta Lake quietly do the heavy lifting behind the scenes. Business impact: dozens of ERP tables moved from raw, multi-language, code-heavy extracts to a single, trusted, human-readable data layer, without adding headcount to the data engineering team. Table of Contents 01  About the Customer 05  Self-Service Data Onboarding 02  The Challenge 06  Medallion Architecture 03  The Solution 07  Business Impact 08  FAQs 09  Conclusion About the Customer Customer Spotlight A Leading Digital Transformation Partner — Europe Our customer is a leading enterprise headquartered in Europe, operating across diverse manufacturing and supply chain divisions. Having already standardized their ERP data through a medallion architecture on Databricks, leadership wanted to go a step further: not only manage ERP data at scale, but also connect it seamlessly into process mining tools to uncover how core processes truly perform in practice. The focus was on gaining operational clarity into workflows such as purchase order management, invoice handling, and procurement cycles. The Challenge Most organizations extracting data from a large ERP system successfully get the data out. The problem isn’t extraction, it’s making that data mean something the moment it lands. Business and IT teams found themselves asking the same questions on repeat: 1Why does the same field appear nine times, with a different value in each column? 2What does “Order Status= 3” actually mean, and who is the source of truth for that mapping? 3How much manual translation and lookup work happens before a single report can be trusted? 4Can business users resolve these issues themselves, without waiting weeks on an IT backlog? 5How do we scale this across dozens of tables without writing dozens of one-off scripts? Two problems came up again and again, and both are far more common across ERP implementations than most leadership teams realize. Multi-Language Columns Because the ERP system was configured for every Order Status the business operates in, a single logical field such as “Item Description” existed as up to nine separate columns, one per language: English, French, German, Spanish, and more. Reports built directly on top of the raw extract had no reliable way of knowing which column to use for which record. In practice, this meant a plant manager in France could open a report and see item names in German, while a sales report for the Spanish market silently pulled blank fields because the Spanish-language column hadn’t been populated for that record. The data was all there; it just wasn’t usable without someone manually deciding, table by table, which language column to trust. Undocumented Domain Values Reference fields like Country, Currency, and Order Status were stored as raw numeric codes rather than readable labels, for example Order Status: 1 = Completed , 2 = In Progress, 3 = Shipped. These mappings lived inside ERP configuration screens, not in the extracted data itself. That meant every downstream report, dashboard, or spreadsheet needed its own copy of the same lookup table, manually kept in sync. When a code changed or a new Order Status was added in the ERP, there was no guarantee every report using it would be updated at the same time, which meant leadership could be looking at the performance chart that was quietly wrong. The Solution Rather than writing custom transformation logic for every table (a solution that ages badly the moment a new table or region gets added), we designed a configuration-driven pipeline built on Azure Databricks, fronted by a self-service web application that puts control directly in the hands of business and functional users. Self-Service Web Portal Business users upload table specifications, review validation results, and queue tables for processing, entirely through a browser. Medallion Architecture Azure Databricks and Delta Lake refine raw extracts through Bronze, Silver, and Gold layers, without table-specific code. Automated Language Resolution Multi-language columns are detected and normalized automatically based on the specification, not hardcoded per table. Centralized Domain Mapping Numeric and coded reference values are resolved against a single, maintained lookup layer instead of scattered spreadsheets. Self-Service Data Onboarding: No Databricks Knowledge Required The centerpiece of the solution is a custom web application that lets a business or functional analyst, not a Databricks engineer, onboard a new ERP table from start to finish. Here’s what that looks like in practice: A business user uploads an Excel-based table specification defining the expected columns, data types, which fields are multi-language, and which fields are domain-coded and how to decode them. The portal validates the specification instantly, flagging missing mandatory columns or mismatches before any data is processed, so problems are caught at the source rather than three reports downstream. Once validation passes, the same user queues the table for processing with a single click. No notebook to open, no cluster to configure, no code to write or review. Behind the scenes, that specification feeds a generic, reusable Databricks framework that already knows how to apply the correct language resolution and domain-value decoding rules, so engineering effort doesn’t scale linearly with the number of tables. In effect, the portal turns “add a new ERP table to the analytics environment” from a data engineering request into a form a finance or operations analyst can complete in minutes, while still enforcing the same rigor and consistency a hand-built pipeline would require. Medallion Architecture on Databricks Once a table is queued through the portal, Azure Databricks takes over: Bronze: Raw ERP extracts are landed as-is, preserving … Continue reading How a Self-Service Data Portal Solved Multi-Language and Domain Value Chaos in ERP Data →

Share Story :

Go Beyond Dashboards- How Databricks Genie Gives Every Business Leader Direct Access to Their Data

Stop Waiting on Reports — Databricks Genie | CloudFronts What You Will Learn Why dashboards alone are no longer enough for fast business decisions What Databricks Genie is and how it enables conversational access to your data How this changes the way finance, sales, and operations teams work What it means for your organization’s AI readiness and long-term decision-making Table of Contents 1. Let’s Start Here 2. The Challenge 3. The Solution — Databricks Genie 4. Business Impact 5. Frequently Asked Questions 6. Conclusion Let’s Start Here Organizations today are not short on data. They have dashboards, reports, and analytics tools in place. But when a business leader needs an answer to a specific question — one that no existing report covers — the usual path is to raise a request, wait for an analyst, and revisit it days later. That delay, small as it seems, adds up. Decisions get deferred. Opportunities get missed. And the data that was meant to drive the business ends up sitting behind a queue. Databricks Genie changes how organizations access their data — by making it conversational. The Challenge Dashboards were built to answer the questions someone thought of in the past. They are excellent for monitoring what is already defined — revenue trends, pipeline stages, operational metrics. But business does not move in straight lines. The moment a leader needs to investigate something outside of what was pre-built, the process breaks down: The question gets raised in a meeting — but no dashboard covers it It gets passed to a data analyst, who adds it to a queue behind other requests Days later, an answer arrives — often too late to influence the decision it was meant to support The result is a quiet, systemic gap between what the business senses and what the data can confirm in time. Leaders fill that gap with instinct. Risks go unspotted. Opportunities pass. Not because the data was not there — but because reaching it took too long. This pattern repeats across every function. Finance cannot investigate a cost anomaly until after month-end close. Sales leadership walks into a quarterly review with numbers someone else prepared. Operations learns about a supplier risk from a weekly report that arrives after the damage is done. The Solution — Databricks Genie Genie is the conversational AI interface built into Azure Databricks. It lets a business leader type a question in plain English — the same way they would ask a colleague — and get an answer drawn from the organization’s actual data, in seconds. There is no form to fill in. No report to request. No specialist to involve for every question. The leader asks, the data responds, and the conversation continues — narrowing, refining, following the next logical question — until the insight is clear enough to act on. The approach rests on three capabilities working together: Conversational access — questions in plain English return precise answers from live data, with no technical skill required from the business user Governed trust — Genie works within existing data permissions; every user sees only what they are authorized to access, and every answer shows the logic behind it Seamless fit — it connects to data the organization already holds, whether from ERP systems, CRM platforms, or operational sources, without requiring a new build This is not a replacement for dashboards. It is what happens between them — the investigative, in-the-moment layer that dashboards were never designed to provide. Business Impact The impact of conversational data access compounds across the organization over time: Decisions get made closer to the moment they matter — leaders investigate anomalies in real time, not after a two-day analysis cycle The right questions finally get asked — when the cost of asking drops to near zero, the volume and quality of insight-driven decisions goes up across every function Data teams focus on higher-value work — instead of fielding one-off requests, analysts build the data models and pipelines that generate lasting value Existing investments go further — Genie extends what the organization has already built, without requiring new infrastructure or a technology overhaul The organization becomes AI-ready — consistent, governed use of data at every level builds the foundation for more advanced AI capabilities to follow The organizations that embrace this shift early will not just be faster. They will be fundamentally better at acting on what they know — and that is an advantage that compounds over time. Frequently Asked Questions Do we need to replace our existing dashboards or BI tools? No. Genie works alongside what you already have. Dashboards remain the right tool for structured, recurring reporting. Genie handles the ad-hoc, investigative questions that dashboards were not built to answer. They complement each other. Does this require technical skills from business users? No. Genie is designed for business users who have no data or SQL background. Questions are asked in plain English — the same way you would ask a colleague — and answers are returned in a readable format without any technical input required. Is the data secure? Can users access data they should not see? Genie inherits the data permissions already configured in your organization’s data environment. Every user sees only what they are already authorized to access. There is no additional access granted by using Genie — governance is built in, not added on. Does our data need to be moved or rebuilt to use Genie? Not necessarily. If your organization’s data — from ERP, CRM, operational systems, or other sources — is already in the Databricks environment, Genie can work with it immediately. For organizations not yet on Databricks, CloudFronts can help assess the right path forward. How is this different from asking an AI chatbot a question about our business? A general AI chatbot answers from its training data — it does not know your organization’s numbers. Genie queries your actual data directly. Every answer is grounded in your real figures, with the source and logic visible, making … Continue reading Go Beyond Dashboards- How Databricks Genie Gives Every Business Leader Direct Access to Their Data →

Share Story :

Designing Metadata-Driven Data Pipelines in Databricks for Scalable Ingestion

Summary In modern data engineering environments, managing ingestion pipelines across multiple source systems becomes increasingly complex as data volume and variety grow. Hardcoded pipelines create maintenance overhead, slow down onboarding of new datasets, and introduce operational risks. This blog explains how a metadata-driven pipeline approach in Databricks can simplify ingestion by using a centralized configuration table to dynamically control pipeline behavior. It highlights how this pattern improves scalability, governance, and maintainability while enabling faster and more reliable data processing. The Real Problem: Hardcoded Pipelines Do Not Scale In many implementations, ingestion pipelines are built separately for each entity or source system. Typical issues include: As the number of entities grows, pipelines become difficult to manage and error-prone. What Is a Metadata-Driven Pipeline? A metadata-driven pipeline shifts control from code to configuration. Instead of writing separate logic for each dataset, we define ingestion behavior in a centralized configuration table. Typical metadata fields include: The pipeline reads this metadata and dynamically executes ingestion logic. Implementation Approach Step 1: Create a Configuration Table A centralized metadata table is created to define ingestion rules. Each row represents one dataset and contains all required configuration. Step 2: Dynamic Pipeline Execution The pipeline reads metadata and loops through each configuration entry. For each entity: No code changes are required when new entities are added. Step 3: Incremental Logic Control Instead of hardcoding: WHERE modifiedon > last_run The incremental field is read from metadata, allowing flexibility across different source systems. Step 4: Integration with Lakehouse Layers Metadata drives ingestion, while Lakehouse layers manage transformation. Why This Approach Works in Enterprise Environments 1. Scalability New entities can be added by inserting a new row in metadata. No pipeline duplication required. 2. Maintainability Changes in incremental logic or source structure are handled centrally. 3. Consistency All pipelines follow the same logic and standards. 4. Governance Metadata provides visibility into: Common Mistakes to Avoid Metadata-driven pipelines require discipline in design. Business Impact Metadata-driven pipelines are not just a technical optimization they are a foundational shift in how data platforms are built and managed. Organizations looking to scale their data engineering capabilities should move away from hardcoded ingestion logic and adopt configuration-driven approaches that support flexibility, governance, and long-term growth. Connect with CloudFronts to get started at transform@cloudfonts.com.

Share Story :

Building a Reliable Bronze Silver Gold Data Pipeline in Databricks for Enterprise Reporting

Summary Modern analytics platforms require structured data pipelines that ensure reliability, consistency, and governance across reporting systems. Traditional ETL approaches often struggle to scale as data volume and complexity increase. This blog explains how the Bronze–Silver–Gold (Medallion) architecture in Databricks provides a scalable and reliable framework for organizing data pipelines. It highlights how each layer serves a specific purpose, enabling better data quality, governance, and seamless integration with reporting tools such as Power BI. The Real Problem: Reporting Pipelines Become Fragile Over Time In many organizations: This leads to unreliable reporting and increased maintenance effort. What Is the Bronze–Silver–Gold Architecture? The Medallion architecture organizes data into three layers: Bronze Layer Raw data ingestion layer. Silver Layer Cleaned and standardized data. Gold Layer Business-ready, reporting-optimized data. Each layer has a clear responsibility. Bronze Layer: Raw Data Ingestion Purpose Key Characteristics Bronze acts as the system of record. Silver Layer: Data Standardization Purpose Key Activities Silver creates reusable datasets across reporting use cases. Gold Layer: Reporting-Ready Data Purpose Key Characteristics Gold tables are consumed directly by reporting tools. Why This Architecture Works 1. Separation of Concerns Each layer has a defined role, reducing complexity. 2. Improved Data Quality Data is progressively refined from raw to curated. 3. Better Performance Reporting queries run on optimized Gold tables. 4. Governance with Unity Catalog Access can be controlled at each layer: Common Implementation Mistakes These mistakes lead to long-term instability. Business Impact To conclude, the Bronze–Silver–Gold architecture provides a strong foundation for building scalable and reliable data pipelines in Databricks. When combined with proper governance and disciplined design, it enables organizations to deliver consistent, high-quality data for analytics and decision-making. We hope you found this article useful. If you would like to explore how AI-powered customer service can improve your support operations, please contact us at transform@cloudfronts.com.

Share Story :

Building a Scalable AI Workforce with Agent Bricks – Part 2

The Challenge of Scaling AI in Enterprises Many organizations invest in AI initiatives but struggle to scale beyond pilot projects. Custom-built solutions are expensive, difficult to govern, and often limited to a single use case. As a result, AI investments fail to deliver sustained business value. Why Automation Alone Is Not Enough Traditional automation relies on rigid rules and predefined workflows. While effective for simple tasks, it cannot adapt to changing business conditions. Enterprises need intelligent systems that can reason, decide, and act autonomously. Understanding AI Agents in Simple Terms AI agents are intelligent software systems that understand goals, plan actions, and execute multi-step workflows with minimal human intervention. Unlike chatbots, AI agents do not just answer questions they act on insights. What Agent Bricks Bring to the Business Agent Bricks are modular, reusable AI agent components that accelerate enterprise AI adoption. They enable organizations to deploy intelligent agents quickly while maintaining security, governance, and compliance. Ask Me Anything: Execution Powered by Agent Bricks In the Ask Me Anything solution, Agent Bricks power the execution layer. They continuously evaluate enterprise data, identify project readiness gaps, and respond to leadership queries in real time. Agent Bricks Workflow Execution (Testing Screenshot) Use Case Spotlight: PMO Assistant at Scale The PMO Assistant built using Agent Bricks operates continuously, monitoring upcoming projects and flagging risks early. This reduces dependency on manual reporting and enables PMOs to focus on proactive delivery management. Business Value of an AI Workforce From a business perspective, Agent Bricks enable faster AI deployment, lower operational costs, and consistent decision-making across departments. Enterprises can scale AI solutions confidently without rebuilding logic for every new use case. Moving from Experiments to Execution To conclude, Agent Bricks help organizations move from isolated AI experiments to production-ready AI solutions. CloudFronts partners with enterprises to build scalable, governed AI workforces that deliver measurable business outcomes. I hope you found this blog useful, and if you would like to discuss anything or explore a future implementation, you can reach out to us at transform@cloudfonts.com.

Share Story :

Advanced Time Travel & Data Recovery Strategies in Delta Lake

In production Databricks environments, data issues such as accidental overwrites, faulty MERGE conditions, or incorrect backfills are common. Delta Lake’s Time Travel is not just a feature – it is a critical recovery and governance mechanism. This blog focuses only on practical recovery strategies that are actually used in real-world production systems. Why Time Travel Is Critical in Production Common failure scenarios include: •a. INSERT OVERWRITE wiping historical data • b. Incorrect MERGE conditions deleting valid records • c. Wrong filters during backfill corrupting data Reprocessing data is expensive and risky. Time Travel enables instant rollback with minimal impact. Version vs Timestamp (What You Should Use) Always prefer version-based time travel for recovery operations. Why version-based recovery is preferred: • a. Precise and deterministic • b. No time zone dependency • c. Safest option for production recovery Use timestamp-based queries only for auditing, not recovery. Identify the Last Safe State Before performing any recovery, always inspect the table history. DESCRIBE HISTORY crm_opportunities; Key fields to review: • a. version • b. timestamp • c. operation • d. userName This history acts as the single source of truth during incidents. Recovery Patterns That Actually Work 1. Partial Data Recovery (Recommended) Recover only the affected records instead of rolling back the entire table. Advantages: • a. No downtime • b. Safe for downstream reports • c. Most production-friendly approach 2. Full Table Restore (Use Carefully) Advantages: •a. Fast and atomic Risks: •a. Impacts all downstream consumers Use this approach only when the entire table is corrupted. Safe Validation Using CLONE Before restoring data in production, validate changes using a clone. Typical use cases: • a. Validate recovered data • b. Compare versions •c. Run business checks Retention & VACUUM (Most Common Mistake) The following command causes permanent data loss: Once vacuumed aggressively, time travel breaks and rollback becomes impossible. Production-Safe Retention Recommended retention: • a. Critical tables: 30 days • b. Reporting tables: 7–14 days Auditing & Root Cause Analysis (RCA) Track who changed data and when: Compare changes between versions: Key Best Practices • a. Capture table version before running risky jobs • b. Always use version-based time travel for recovery • c. Prefer partial recovery over full restores • d. Avoid aggressive VACUUM operations • e. Extend retention for critical tables • f. Validate using CLONE before restoring To conclude, Delta Lake Time Travel is not a backup mechanism, but it is the fastest and safest recovery tool in Databricks. When used correctly, it prevents downtime, reduces reprocessing cost, and improves production reliability. For enterprise Databricks pipelines, mastering this capability is mandatory, not optional. We hope you found this blog useful, and if you would like to discuss anything, you can reach out to us at transform@cloudfronts.com

Share Story :

SEARCH BLOGS:

FOLLOW CLOUDFRONTS BLOG :


Categories

Secured By miniOrange