<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Data &amp; Cloud Migration Archives - Inferenz</title>
	<atom:link href="https://inferenz.ai/category/data-cloud-migration/feed/" rel="self" type="application/rss+xml" />
	<link></link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 07:08:45 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	

<image>
	<url>https://inferenz.ai/wp-content/uploads/2026/01/freepik__the-logo-rotates-in-place__11558-1.png</url>
	<title>Data &amp; Cloud Migration Archives - Inferenz</title>
	<link></link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Building Audit-Ready Data Lineage That Survives a CMS Review</title>
		<link>https://inferenz.ai/blogs/building-audit-ready-data-lineage-that-survives-a-cms-review/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 07:07:58 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<category><![CDATA[Healthcare]]></category>
		<category><![CDATA[Hospice and Palliative]]></category>
		<guid isPermaLink="false">https://inferenz.ai/blogs//</guid>

					<description><![CDATA[<p>Most hospices can answer that question eventually. Very few can answer it inside the 45 calendar days a Medicare Administrative Contractor (MAC) allows for an Additional Documentation Request, and fewer still can answer it the same way twice if a second reviewer asks.</p>
<p>The post <a href="https://inferenz.ai/blogs/building-audit-ready-data-lineage-that-survives-a-cms-review/">Building Audit-Ready Data Lineage That Survives a CMS Review</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h2>Summary</h2>
<p><em>A Medicare Administrative Contractor&#8217;s additional documentation request gives a hospice 45 calendar days to produce a complete record for the one claim it&#8217;s questioning. Most compliance teams can pull that claim in minutes. Far fewer can show, system by system, exactly how it got there, and that second skill is what a real audit tests.</em></p>
<h2>Introduction</h2>
<p>The letter rarely looks like the start of something serious. It names a patient, a date of service, a benefit period, and a deadline. From there, someone at the hospice has to rebuild the full story behind that record: which system captured the physician&#8217;s narrative, when the election statement was signed, whether the diagnosis on the claim still matches the diagnosis in the chart, and who touched the file along the way. That rebuild is exactly what <a href="https://inferenz.ai/services/data-quality-governance-and-compliance/">audit-ready data lineage</a> is built to answer, ideally long before any letter shows up.</p>
<p>Most hospices can answer that question eventually. Very few can answer it inside the 45 calendar days a Medicare Administrative Contractor (MAC) allows for an <a href="https://www.cms.gov/data-research/monitoring-programs/medicare-fee-service-compliance-programs/medical-review-education/additional-documentation-request">Additional Documentation Request</a>, and fewer still can answer it the same way twice if a second reviewer asks. That gap, between having the data and being able to prove where it came from, is the specific problem data lineage is built to close. CMS and the HHS Office of Inspector General have both gotten sharper at testing for exactly that gap.</p>
<h2>What &#8220;Audit-ready&#8221; actually means, and why it’s different from &#8220;Data-ready&#8221;</h2>
<p>A hospice can be data-ready and still fail an audit. Data-ready means the EMR holds a diagnosis, the billing system holds a claim, and the two roughly agree. Audit-ready means the hospice can show, for any single record an auditor picks, the full chain of custody: which system created the data element, which system or person changed it, when the change happened, and why. That chain is data lineage, and it&#8217;s a meaningfully different deliverable from data quality alone.</p>
<p>The distinction matters because CMS and OIG reviewers aren&#8217;t just checking whether a number is correct. They&#8217;re checking whether the <a href="https://inferenz.ai/industries/healthcare/hospice-and-palliative-care/">hospice care agencies</a> can reconstruct, on demand, how that number came to be. A hospice that connects its EMR, billing, and referral intake data has usually solved the double-entry problem that eats staff time. It still has separate work to do on the lineage problem, because a clean interface moves data between systems without automatically recording where each value originated or who last touched it.</p>
<h2>The audits that actually show up at a hospice&#8217;s door</h2>
<p>Three oversight mechanisms account for almost every real compliance event a hospice will face, and each one tests data lineage in a slightly different way.</p>
<ul>
<li><strong>OIG compliance audits, built on statistical sampling and extrapolation.</strong> The HHS Office of Inspector General runs an ongoing series of hospice compliance audits, and it doesn&#8217;t review every claim to do it. It pulls a sample, commonly around 100 claims, checks each one against Medicare requirements, and when the error rate is high enough, extrapolates the dollar impact across the full population of claims the hospice billed for that period. CMS&#8217;s own 2026 oversight update pointed to how real that exposure has become, noting that enhanced review in four states alone had already produced <a href="https://www.cms.gov/newsroom/press-releases/cms-proposes-new-transparency-measures-strengthen-oversight-hospice-providers">more than 200 hospice Medicare enrollment revocations</a>. OIG picks which hospices to review in the first place using computer matching, data mining, and data analysis techniques run against claims data, so audit risk starts building long before any letter shows up. A hospice that can&#8217;t trace a sampled claim back through its clinical and billing history has no real way to challenge either the finding or the extrapolation that follows it.</li>
<li><strong>MAC medical review, through Additional Documentation Requests and Targeted Probe and Educate.</strong> MACs run both broad and targeted reviews of hospice claims. Under Targeted Probe and Educate, a MAC pulls 20 to 40 claims per round from providers with the highest denial rates or the most unusual billing patterns, running up to three rounds with individualized education between each. The clock on any ADR starts the moment it lands.</li>
<li><strong>HQRP compliance, enforced through the HOPE submission threshold.</strong> Since October 1, 2025, the Hospice Outcomes and Patient Evaluation</li>
<li>(HOPE) tool has replaced the Hospice Item Set, and every record now has to move through iQIES, the only channel CMS accepts since the legacy QIES/ASAP system was retired. CMS requires 90% of HOPE records submitted within 30 days of the relevant admission, discharge, or update-visit date, or the hospice loses 4 percentage points off its Annual Payment Update. That threshold functions as a lineage problem more than a paperwork one: the data has to flow cleanly from the point of clinical assessment into iQIES without someone manually re-keying it under deadline pressure.</li>
</ul>
<p>CMS has also been raising the general level of scrutiny hospices operate under. In 2026, the agency <a href="https://www.cms.gov/newsroom/press-releases/cms-proposes-new-transparency-measures-strengthen-oversight-hospice-providers">proposed new transparency measures</a> built partly on a Service and Spending Variation Index that scores hospices using claims-based utilization metrics, publishes those scores, and flags high-scoring hospices for additional review. The same announcement noted that roughly <a href="https://www.cms.gov/newsroom/press-releases/cms-proposes-new-transparency-measures-strengthen-oversight-hospice-providers">20% of hospices were out of compliance with HQRP reporting requirements</a> in CY 2025, a rate CMS described as comparable to prior years. That consistency is the tell and this is a widespread industry gap.</p>
<h2>What CMS&#8217;s Medicare Advantage Audit Framework Signals for Hospice</h2>
<p>Hospice program audits don&#8217;t run on the same protocol as Medicare Advantage program audits, and it&#8217;s worth being precise about that before drawing any comparison. Still, the logic CMS applies to one Medicare program tends to surface in the others eventually, and the classification system CMS already uses for Part C and Part D audits is a useful preview of where hospice oversight is headed.</p>
<p>CMS&#8217;s Part C and Part D program audit framework sorts findings into an Observation, for noncompliance that doesn&#8217;t need a formal fix, a Corrective Action Required (CAR), for noncompliance that does, and an Invalid Data Submission (IDS) finding, reserved for cases where a sponsor cannot produce an accurate, complete “universe” of records and CMS cannot determine compliance as a result. CMS has also been evaluating Compliance Program Effectiveness (CPE) less as a checklist of written policies and more as a live conversation about how a plan actually detects and corrects problems as they happen.</p>
<p>The IDS classification is really a test of whether a sponsor can produce a defensible, complete record on demand, and a single incorrect answer is a much smaller problem than a universe of records nobody can reconstruct. Risk Adjustment Data Validation (RADV) audits, the mechanism CMS uses to verify Medicare Advantage risk-adjustment payments, run on a related logic: the diagnosis on a claim has to trace back to a specific, contemporary clinical encounter, or CMS claws back the payment.</p>
<p>Hospice oversight runs on a different vocabulary altogether: OIG extrapolation, MAC documentation requests, and SSVI scoring. The underlying test is the same one regardless. Can the organization produce a complete, source-traceable universe of records on demand? A <a href="https://inferenz.ai/healthcare-solutions/mpi-and-patient-360/">governed data lineage layer</a> exists to pass exactly that test.</p>
<p><a href="https://inferenz.ai/blogs/how-hospice-organizations-build-a-cms-ready-data-foundation/"><img fetchpriority="high" decoding="async" class="alignnone size-full wp-image-16934" src="https://inferenz.ai/wp-content/uploads/2026/10/See-how-hospice-organizations-can-build-a-CMS-ready-data-foundation-efficiently.jpg" alt="See how hospice organizations can build a CMS-ready data foundation efficiently" width="1340" height="350" srcset="https://inferenz.ai/wp-content/uploads/2026/10/See-how-hospice-organizations-can-build-a-CMS-ready-data-foundation-efficiently.jpg 1340w, https://inferenz.ai/wp-content/uploads/2026/10/See-how-hospice-organizations-can-build-a-CMS-ready-data-foundation-efficiently-300x78.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/10/See-how-hospice-organizations-can-build-a-CMS-ready-data-foundation-efficiently-1024x267.jpg 1024w, https://inferenz.ai/wp-content/uploads/2026/10/See-how-hospice-organizations-can-build-a-CMS-ready-data-foundation-efficiently-768x201.jpg 768w" sizes="(max-width: 1340px) 100vw, 1340px" /></a></p>
<h2>Data Lineage vs. Data Provenance: A Distinction That Matters Once CMS Asks</h2>
<p>The two terms get used interchangeably, and in a hospice audit context, the difference is worth keeping straight. Data provenance answers where a piece of data originated: which system, which form, which staff member entered it first. It&#8217;s answers the fuller question: every system and transformation that record passed through afterward, including edits, merges, and exports, on its way to the CMS submission.</p>
<h3><strong>Data Provenance vs. Data Lineage</strong></h3>
<table>
<tbody>
<tr>
<td></td>
<td><strong>Data Provenance</strong></td>
<td><strong>Data Lineage</strong></td>
</tr>
<tr>
<td><strong>What it answers</strong></td>
<td>Where a piece of data originated</td>
<td>Everywhere that data went afterward</td>
</tr>
<tr>
<td><strong>Scope</strong></td>
<td>A single point in time: the system, form, or staff member who first entered it</td>
<td>The full path: every system, edit, merge, and export the record passed through</td>
</tr>
<tr>
<td><strong>Example (election statement)</strong></td>
<td>Confirms the election was signed in the EMR on a given date</td>
<td>Confirms that update also reached billing and the referral record before any revocation, benefit-period change, or new admission touched the same patient again</td>
</tr>
<tr>
<td><strong>What it catches</strong></td>
<td>Who entered a value, and when</td>
<td>Desynchronization between systems, the exact pattern CMS&#8217;s hospital-hospice claims edits are built to flag</td>
</tr>
<tr>
<td><strong>Why a hospice needs both</strong></td>
<td>Establishes the starting point of the record</td>
<td>Establishes whether that record stayed consistent all the way to the CMS submission</td>
</tr>
</tbody>
</table>
<p>A hospice election statement is a good example of why the difference matters. Provenance tells you the election was signed in the EMR on a given date. Lineage tells you whether that election status update also reached billing and the referral record before a revocation, a benefit-period change, or a new admission touched the same patient again. CMS&#8217;s newer claims edits, the ones comparing hospital and hospice claims for overlapping services and flagging admission-date and billing-date mismatches, are built to catch exactly the kind of desynchronization a gap like this creates.</p>
<h2>What a Defensible Audit Trail Actually Has to Show</h2>
<p><img decoding="async" class="alignnone size-full wp-image-16936" src="https://inferenz.ai/wp-content/uploads/2026/10/What-a-Defensible-Audit-Trail-Actually-Has-to-Show.png" alt="What a Defensible Audit Trail Actually Has to Show" width="1804" height="872" srcset="https://inferenz.ai/wp-content/uploads/2026/10/What-a-Defensible-Audit-Trail-Actually-Has-to-Show.png 1804w, https://inferenz.ai/wp-content/uploads/2026/10/What-a-Defensible-Audit-Trail-Actually-Has-to-Show-300x145.png 300w, https://inferenz.ai/wp-content/uploads/2026/10/What-a-Defensible-Audit-Trail-Actually-Has-to-Show-1024x495.png 1024w, https://inferenz.ai/wp-content/uploads/2026/10/What-a-Defensible-Audit-Trail-Actually-Has-to-Show-768x371.png 768w, https://inferenz.ai/wp-content/uploads/2026/10/What-a-Defensible-Audit-Trail-Actually-Has-to-Show-1536x742.png 1536w" sizes="(max-width: 1804px) 100vw, 1804px" /></p>
<p>When a MAC, a state surveyor, or an OIG auditor asks a hospice to defend a record, they&#8217;re rarely asking for a single data point. They&#8217;re asking for a reconstruction. A defensible audit trail needs to show, for any patient and any date range:</p>
<ul>
<li><strong>Source-to-submission mapping.</strong> Which system captured the original clinical assessment, referral, or diagnosis, and every hop that data made on its way into a HOPE record, a claim, or an iQIES submission.</li>
<li><strong>Claims-to-clinical reconciliation.</strong> Whether the level of care, visit dates, and diagnosis codes on the claim still match the clinical documentation behind them, since a mismatch here is the specific pattern CMS&#8217;s newer hospital-hospice billing edits are designed to catch.</li>
<li><strong>Election and revocation history.</strong> A timestamped, cross-system record of every election statement, revocation, and benefit-period change, so EMR, billing, and referral records can&#8217;t disagree about a patient&#8217;s current status.</li>
<li><strong>Grievance and appeals traceability.</strong> A hospice&#8217;s own complaint and grievance records, plus any Medicare claim appeals filed on a beneficiary&#8217;s behalf, need the same source-to-outcome trail as clinical data, because surveyors and MACs both ask for it.</li>
<li><strong>AI model lineage, where applicable.</strong> If a hospice uses AI to support documentation, eligibility screening, or care planning, an auditor can reasonably ask which data informed that output and what a clinician reviewed before it reached the record. Most hospices haven&#8217;t built for this yet, and it&#8217;s becoming a standing expectation.</li>
</ul>
<h2>Where Hospice Data Lineage Actually Breaks</h2>
<p>Four gaps account for most of the lineage failures we see in hospice organizations, and every one of them is fixable without touching the underlying systems.</p>
<ol>
<li>The first is the referral-to-EMR handoff, where a referral arriving by fax or portal gets keyed in as a new patient record and later turns out to duplicate one created through a different intake path.</li>
<li>The second is election status drift, where a revocation updates in one system only, creating the exact mismatch CMS&#8217;s audit logic is built to flag.</li>
<li>The third is the pharmacy and durable medical equipment gap, where medication and equipment data still move by phone and fax while everything else has been automated, leaving a hole in the same clinical record CMS is checking.</li>
<li>The fourth, and increasingly the most consequential, is the spreadsheet workaround: the export someone pulls into Excel to reconcile two systems by hand, which quietly becomes the record of truth nobody can trace back to its source once an auditor asks where a number came from.</li>
</ol>
<p>These line breaks mean that hospice care agencies will repeatedly bank on erroneous data for their processes. Fixing the data lineage is crucial for these organizations to rely on accurate data and make informed decisions.</p>
<h2>Building the Lineage Layer Without a Rip-and-Replace</h2>
<p><img decoding="async" class="alignnone size-full wp-image-16935" src="https://inferenz.ai/wp-content/uploads/2026/10/Building-the-Lineage-Layer-Without-a-Rip-and-Replace.png" alt="Building the Lineage Layer Without a Rip-and-Replace" width="1804" height="872" srcset="https://inferenz.ai/wp-content/uploads/2026/10/Building-the-Lineage-Layer-Without-a-Rip-and-Replace.png 1804w, https://inferenz.ai/wp-content/uploads/2026/10/Building-the-Lineage-Layer-Without-a-Rip-and-Replace-300x145.png 300w, https://inferenz.ai/wp-content/uploads/2026/10/Building-the-Lineage-Layer-Without-a-Rip-and-Replace-1024x495.png 1024w, https://inferenz.ai/wp-content/uploads/2026/10/Building-the-Lineage-Layer-Without-a-Rip-and-Replace-768x371.png 768w, https://inferenz.ai/wp-content/uploads/2026/10/Building-the-Lineage-Layer-Without-a-Rip-and-Replace-1536x742.png 1536w" sizes="(max-width: 1804px) 100vw, 1804px" /></p>
<p>Fixing this leakage starts with a governed data layer that sits underneath the EMR, the billing platform, and the referral tools a hospice&#8217;s staff already know and use every day. That layer tracks, at the record level, where data entered the organization, every system it moved through, and every transformation applied along the way. In practice, that means:</p>
<ul>
<li><strong>Mapping the systems first, before the tools.</strong> Before evaluating any data lineage software, a hospice needs a clear map of every system that touches a patient record: EMR, billing, referral intake, pharmacy, DME, and any AI tools layered on top. Lineage mapping only works when it starts from a complete inventory.</li>
<li><strong>Treating lineage as one feature of a broader governance foundation.</strong> Standalone data lineage tracking works best as part of a wider data quality management and governance framework. A point tool that only visualizes lineage after the fact does little to fix the upstream data quality gaps that create the mismatches auditors flag in the first place.</li>
<li><strong>Building automated, near-real-time interfaces.</strong> The same bidirectional integration principles that solve hospice EMR-to-billing double entry also generate a cleaner lineage record along the way, because a system that writes and reads data automatically leaves a far more traceable trail than one staff bridge by hand.</li>
<li><strong>Treating AI documentation and eligibility tools as lineage sources.</strong> Any AI system touching clinical documentation, <a href="https://inferenz.ai/healthcare-solutions/caregence-agents/hospice-eligibility-agent/">eligibility screening</a>, or care planning needs its own record of what data it used and what a clinician confirmed, a gap <a href="https://inferenz.ai/blogs/why-most-hospice-ai-projects-fail-without-data-readiness/">Inferenz has covered separately</a> in the context of hospice AI readiness generally.</li>
</ul>
<p>A hospice&#8217;s own <a href="https://inferenz.ai/services/data-quality-governance-and-compliance/">data quality, governance, and compliance program</a> is the natural home for this work, since lineage without governance mostly produces a very detailed record of an ungoverned mess. Inferenz has built this kind of foundation for <a href="https://inferenz.ai/case-studies/building-an-enterprise-data-platform-from-the-ground-up-for-a-post-acute-care-organisation/">post-acute care organizations moving off 32 disconnected source systems onto one governed data platform</a>, the same underlying problem most hospices face at a smaller scale.</p>
<h2>What to Have Ready Before the Letter Arrives</h2>
<p>Continuous audit readiness looks different from the once-a-year scramble most hospices still run today. The organizations that stop dreading MAC letters and OIG notices tend to share a few habits: they can pull a complete case file, source-to-submission, for any patient within minutes; their compliance officer reviews a sample of records against CMS&#8217;s current claims-edit logic on a standing schedule; and their <a href="https://inferenz.ai/industries/healthcare/hospice-and-palliative-care/">hospice and palliative care</a> operations team treats data governance as an ongoing operational function.</p>
<p>That shift, from periodic to continuous, is also where CMS&#8217;s own compliance language across every Medicare program keeps pointing: fewer point-in-time checklists, more real-time evidence that an organization catches and corrects its own problems before an outside reviewer does it.</p>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-16933" src="https://inferenz.ai/wp-content/uploads/2026/10/Ready-to-know-more-about-building-an-audit-ready-data-foundation-for-your-hospice.jpg" alt="Ready to know more about building an audit-ready data foundation for your hospice?" width="1340" height="350" srcset="https://inferenz.ai/wp-content/uploads/2026/10/Ready-to-know-more-about-building-an-audit-ready-data-foundation-for-your-hospice.jpg 1340w, https://inferenz.ai/wp-content/uploads/2026/10/Ready-to-know-more-about-building-an-audit-ready-data-foundation-for-your-hospice-300x78.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/10/Ready-to-know-more-about-building-an-audit-ready-data-foundation-for-your-hospice-1024x267.jpg 1024w, https://inferenz.ai/wp-content/uploads/2026/10/Ready-to-know-more-about-building-an-audit-ready-data-foundation-for-your-hospice-768x201.jpg 768w" sizes="auto, (max-width: 1340px) 100vw, 1340px" /></a></p>
<h2 style="margin: 16.0pt 0cm 10.0pt 0cm;">Frequently Asked Questions</h2>
<p>The post <a href="https://inferenz.ai/blogs/building-audit-ready-data-lineage-that-survives-a-cms-review/">Building Audit-Ready Data Lineage That Survives a CMS Review</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Databricks Genie One: Inside the Agentic Coworker Turning Business Data into Action</title>
		<link>https://inferenz.ai/blogs/databricks-genie-one-inside-the-agentic-coworker-turning-business-data-into-action/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Fri, 14 Aug 2026 07:41:16 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://inferenz.ai/blogs//</guid>

					<description><![CDATA[<p>Databricks Genie One is the agentic coworker on Databricks' Data Intelligence Platform, letting any business user, not just analysts, ask questions of governed data, save repeatable skills, and automate recurring work through scheduled tasks.</p>
<p>The post <a href="https://inferenz.ai/blogs/databricks-genie-one-inside-the-agentic-coworker-turning-business-data-into-action/">Databricks Genie One: Inside the Agentic Coworker Turning Business Data into Action</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h2><b><span data-contrast="none">Summary</span></b><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></h2>
<p><b><i><span data-contrast="none">Databricks Genie One</span></i></b><i><span data-contrast="none"> is the agentic coworker on Databricks&#8217; Data Intelligence Platform, letting any business user, not just analysts, ask questions of governed data, save repeatable skills, and automate recurring work through scheduled tasks. It runs on </span></i><b><i><span data-contrast="none">Genie Agents</span></i></b><i><span data-contrast="none"> and </span></i><b><i><span data-contrast="none">Metric Views</span></i></b><i><span data-contrast="none"> for descriptive analytics, then extends into predictive use cases, like hospital readmission risk, through custom agents built on the </span></i><b><i><span data-contrast="none">Mosaic AI Agent Framework</span></i></b><i><span data-contrast="none">. This guide covers setup, accuracy and cost tuning, and field-tested best practices for rolling Genie One out across marketing, finance, sales, HR, and clinical operations teams.</span></i><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><img loading="lazy" decoding="async" class="alignnone size-full wp-image-16601" src="https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-Architecture.png" alt="Databricks Genie One: End-to-End Architecture" width="1340" height="890" srcset="https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-Architecture.png 1340w, https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-Architecture-300x199.png 300w, https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-Architecture-1024x680.png 1024w, https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-Architecture-768x510.png 768w" sizes="auto, (max-width: 1340px) 100vw, 1340px" /></p>
<p><span data-contrast="none">A care coordination lead at a mid-size hospital system used to lose two days to a single readmission report: pulling numbers from three dashboards, emailing an analyst, and hoping the definitions matched. That wait is going away.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">Databricks Genie One, the agentic coworker built into the Data Intelligence Platform, lets any business user, not just analysts, ask questions of governed data, teach it repeatable skills, and hand it recurring work through scheduled tasks. </span><a href="https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025"><span data-contrast="none">Gartner expects 40% of enterprise applications to embed task-specific AI agents by the end of 2026, up from un</span></a><span data-contrast="none">der 5% in 2025, and Genie One is Databricks&#8217; clearest answer to that shift yet.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">This guide breaks down what Genie One does, the Genie Agents and Metric Views it runs on, how teams extend it into predictive work, and where the real accuracy and cost trade-offs live.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><img loading="lazy" decoding="async" class="alignnone size-full wp-image-16602" src="https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One.png" alt="care coordination lead at a mid-size hospital system used to lose two days to a single readmission report." width="1340" height="907" srcset="https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One.png 1340w, https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-300x203.png 300w, https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-1024x693.png 1024w, https://inferenz.ai/wp-content/uploads/2026/08/Databricks-Genie-One-768x520.png 768w" sizes="auto, (max-width: 1340px) 100vw, 1340px" /></p>
<h2><span class="TextRun SCXW18352622 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW18352622 BCX8" data-ccp-parastyle="heading 2">What Is Databricks Genie One?</span></span></h2>
<p><span data-contrast="none">Databricks Genie One is the general-purpose chat surface of the Data Intelligence Platform, built for people who have never written a line of SQL. </span></p>
<p><span data-contrast="none">Where a <strong>Genie Agent</strong> is a governed, domain-scoped data building block, Genie One is the coworker any business user talks to. It draws its verified context from <strong>Genie Ontology</strong> and its trusted data and metrics from one or more <strong>Genie Agents</strong>, then layers on the capabilities a conversational agent needs to actually get work done. It answers questions, drafts documents and artifacts, takes action through MCP tools, etc. The two capabilities most relevant to day-to-day adoption are covered in depth below: <strong>skills</strong> and <strong>scheduled tasks</strong>.</span></p>
<ul>
<li><strong>Self-service in plain language: </strong>any business team can ask a question of governed data in natural language and get an answer without learning a BI tool or waiting on an analyst.</li>
<li><strong>Answers grounded in your data: </strong>accuracy comes from Genie Ontology, Genie Agents, and Metric Views, which give Genie One a working understanding of the organization&#8217;s own vocabulary, tables, and certified business metrics rather than a generic model guess.</li>
<li><strong>Governed by design: </strong>every answer, regardless of which team asks or how the question is phrased, is scoped to the same Unity Catalog permissions and lineage that govern the underlying data, so results stay compliant with existing access and governance policies.</li>
</ul>
<p><a href="https://www.databricks.com/company/newsroom/press-releases/databricks-launches-genie-one-all-new-agentic-coworker-every-team"><span data-contrast="none">Databricks announced it on June 16, 2026 at the Data + AI Summit</span></a><span data-contrast="none">, positioning it as a real step up from the original Genie, which only answered questions about data already sitting inside Databricks.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">Genie One reaches further. It works across structured and unstructured data, inside and outside the platform, and it ships on web, iOS, and Android. Marketing, finance, sales, HR, and clinical operations teams all get the same coworker, grounded in the same Unity Catalog permissions and lineage that govern everything else on the platform.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">At </span>Inferenz<span data-contrast="none">, a <a href="https://inferenz.ai/">data and AI solutions-led services company</a> and Databricks partner, our team has believed that the differentiators sit underneath: Genie Agents, Metric Views, and, once the question turns predictive, </span><a href="https://inferenz.ai/services/generative-and-agentic-ai/"><span data-contrast="none">custom agents, built on the Mosaic AI Agent Framework</span></a><span data-contrast="none">.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<h2><span class="TextRun SCXW137060867 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW137060867 BCX8" data-ccp-parastyle="heading 2">Skills and scheduled tasks: how Genie One learns to work like you do</span></span><span class="EOP Selected SCXW137060867 BCX8" data-ccp-props="{&quot;335559738&quot;:340,&quot;335559739&quot;:140,&quot;335572079&quot;:4,&quot;335572080&quot;:4,&quot;335572081&quot;:13684944,&quot;469789806&quot;:&quot;single&quot;}"> </span></h2>
<p><span data-contrast="none">Two features separate Genie One from a chatbot that forgets everything after each session.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">A skill is a task you teach the coworker once and reuse by name. Ask it to build your weekly metrics report, and it saves that approach as reusable, inspectable text under the open Agent Skills standard, the same convention Genie Code runs on. User skills currently sit in Public Preview, saved privately to your workspace, and Genie One applies one automatically unless you @-mention it directly.</span></p>
<p><img loading="lazy" decoding="async" class="alignnone size-full wp-image-16604" src="https://inferenz.ai/wp-content/uploads/2026/08/Skills-and-scheduled-tasks-how-Genie-One-learns-to-work-like-you-do.png" alt="Skills and scheduled tasks: how Genie One learns to work like you do " width="1340" height="799" srcset="https://inferenz.ai/wp-content/uploads/2026/08/Skills-and-scheduled-tasks-how-Genie-One-learns-to-work-like-you-do.png 1340w, https://inferenz.ai/wp-content/uploads/2026/08/Skills-and-scheduled-tasks-how-Genie-One-learns-to-work-like-you-do-300x179.png 300w, https://inferenz.ai/wp-content/uploads/2026/08/Skills-and-scheduled-tasks-how-Genie-One-learns-to-work-like-you-do-1024x611.png 1024w, https://inferenz.ai/wp-content/uploads/2026/08/Skills-and-scheduled-tasks-how-Genie-One-learns-to-work-like-you-do-768x458.png 768w" sizes="auto, (max-width: 1340px) 100vw, 1340px" /></p>
<p><span data-contrast="none">A scheduled task runs that same logic on a cadence you set, in plain English (something like &#8220;send me a daily briefing of new customer reviews&#8221;), then posts results into a chat thread plus an email. Schedules currently cap at daily frequency by default and need the Databricks SQL access entitlement to create or run.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<table data-tablestyle="MsoNormalTable" data-tablelook="0" aria-rowcount="6" aria-colcount="3">
<tbody>
<tr aria-rowindex="1">
<td data-celllook="69905"><b><span data-contrast="none">Dimension</span></b><span data-ccp-props="{}"> </span></td>
<td data-celllook="69905"><b><span data-contrast="none">Skills</span></b><span data-ccp-props="{}"> </span></td>
<td data-celllook="69905"><b><span data-contrast="none">Scheduled Tasks</span></b><span data-ccp-props="{}"> </span></td>
</tr>
<tr aria-rowindex="2">
<td data-celllook="69905"><span data-contrast="none">Trigger</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Manual, or auto-detected by Genie One</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Runs automatically on a set cadence</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="3">
<td data-celllook="69905"><span data-contrast="none">Best for</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Repeatable, on-demand tasks</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Recurring reports and monitoring</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="4">
<td data-celllook="69905"><span data-contrast="none">Output</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Chat response, document, or action</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Chat message plus an email notification</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="5">
<td data-celllook="69905"><span data-contrast="none">How you set it up</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Ask Genie One to save an approach, or build one directly</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Natural-language request, or a manual form: Title, Instructions, Connections, Schedule, Timezone</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="6">
<td data-celllook="69905"><span data-contrast="none">Governance note</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Ordinary files in your workspace folder, not hidden settings</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Capped at daily frequency by default; needs the SQL access entitlement</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
</tbody>
</table>
<h2><a href="https://inferenz.ai/healthcare-solutions/caregence-agents/"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-16612" src="https://inferenz.ai/wp-content/uploads/2026/08/Caregence-pairs-Genie-style-conversational-agents-for-hospital-home-care-and-hospice-operators-1.jpg" alt="Caregence-pairs-Genie-style-conversational-agents-for-hospital,-home,-care-and-hospice-operators" width="1340" height="350" srcset="https://inferenz.ai/wp-content/uploads/2026/08/Caregence-pairs-Genie-style-conversational-agents-for-hospital-home-care-and-hospice-operators-1.jpg 1340w, https://inferenz.ai/wp-content/uploads/2026/08/Caregence-pairs-Genie-style-conversational-agents-for-hospital-home-care-and-hospice-operators-1-300x78.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/08/Caregence-pairs-Genie-style-conversational-agents-for-hospital-home-care-and-hospice-operators-1-1024x267.jpg 1024w, https://inferenz.ai/wp-content/uploads/2026/08/Caregence-pairs-Genie-style-conversational-agents-for-hospital-home-care-and-hospice-operators-1-768x201.jpg 768w" sizes="auto, (max-width: 1340px) 100vw, 1340px" /></a><span class="TextRun SCXW166749828 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW166749828 BCX8" data-ccp-parastyle="heading 2">Genie Agents: What They Are, and How to Set One Up</span></span></h2>
<p><span class="TextRun SCXW140801494 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="auto"><span class="NormalTextRun SCXW140801494 BCX8">None of this works without trustworthy data underneath it, and </span><span class="NormalTextRun SCXW140801494 BCX8">that&#8217;s</span><span class="NormalTextRun SCXW140801494 BCX8"> the job of a </span></span><span class="TextRun SCXW140801494 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="auto"><span class="NormalTextRun SCXW140801494 BCX8">Genie Agent</span></span><span class="TextRun SCXW140801494 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="auto"><span class="NormalTextRun SCXW140801494 BCX8"> (what Databricks called a Genie Space through mid-2026). It&#8217;s a curated, conversational layer built over roughly 30 tables or views at most in current releases, configured with three things: instructions that teach Genie your vocabulary, SQL examples that anchor its query generation, and trusted assets, certified metrics Genie reuses instead of regenerating from scratch.</span></span><span class="EOP Selected SCXW140801494 BCX8" data-ccp-props="{&quot;335559739&quot;:160}"> </span></p>
<p><img loading="lazy" decoding="async" class="alignnone size-full wp-image-16605" src="https://inferenz.ai/wp-content/uploads/2026/08/Setting-up-Genie-Agent-step-by-step-workflow.png" alt="Setting up Genie Agent step-by-step workflow" width="1340" height="661" srcset="https://inferenz.ai/wp-content/uploads/2026/08/Setting-up-Genie-Agent-step-by-step-workflow.png 1340w, https://inferenz.ai/wp-content/uploads/2026/08/Setting-up-Genie-Agent-step-by-step-workflow-300x148.png 300w, https://inferenz.ai/wp-content/uploads/2026/08/Setting-up-Genie-Agent-step-by-step-workflow-1024x505.png 1024w, https://inferenz.ai/wp-content/uploads/2026/08/Setting-up-Genie-Agent-step-by-step-workflow-768x379.png 768w" sizes="auto, (max-width: 1340px) 100vw, 1340px" /></p>
<p><span data-contrast="auto">A user&#8217;s question becomes SQL, runs on a SQL Warehouse, and comes back with the generated query attached for verification. Nothing here is a black box. As of mid-2026, Agent Mode adds iterative reasoning for open-ended &#8220;why&#8221; and &#8220;what-if&#8221; questions, running several queries and returning a cited report instead of a single number.</span><span data-ccp-props="{&quot;335559739&quot;:160}"> </span></p>
<p><span data-contrast="auto">Setting one up is mostly configuration, not code:</span><span data-ccp-props="{&quot;335559739&quot;:160}"> </span></p>
<ol>
<li><strong>Prerequisites</strong><span data-contrast="auto">: a Pro or Serverless SQL Warehouse, Unity Catalog SELECT privileges, and well-documented tables (column comments, keys, certified tags).</span><span data-ccp-props="{&quot;335559739&quot;:160}"><br />
</span></li>
<li><b><span data-contrast="auto">Create the agent</span></b><span data-contrast="auto">: pick your Unity Catalog sources, name it, attach a warehouse.</span><span data-ccp-props="{&quot;335559739&quot;:160}"><br />
</span></li>
<li><b><span data-contrast="auto">Add data assets</span></b><span data-contrast="auto">: start with 5 to 15 tables or Metric Views in one business domain, not the whole warehouse.</span></li>
<li><b><span data-contrast="auto">Add context</span></b><span data-contrast="auto">: instructions, SQL examples, trusted assets. This is the single highest-leverage step for accuracy.</span></li>
<li><b><span data-contrast="auto">Add sample questions</span></b><span data-contrast="auto">, and enable Agent Mode if users will ask open-ended, multi-step questions.</span></li>
<li><b><span data-contrast="auto">Test, benchmark, and publish</span></b><span data-contrast="auto">, then connect it to Genie One through Unity Catalog groups.</span><span data-ccp-props="{&quot;335559739&quot;:160}"> </span></li>
</ol>
<p><span data-contrast="auto">At Inferenz, this is exactly the discipline we bring to </span><a href="https://inferenz.ai/blogs/databricks-unity-catalog-building-a-unified-data-governance-layer-in-modern-data-platforms/"><span data-contrast="none">Unity Catalog rollouts across healthcare Lakehouse environments</span></a><span data-contrast="auto">: narrow scope first, governance built in from step one, not bolted on after.</span><span data-ccp-props="{&quot;335559739&quot;:160}"> </span></p>
<h2><span class="TextRun SCXW33688343 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW33688343 BCX8" data-ccp-parastyle="heading 2">Metric views: why Genie One never gives two different answers</span></span><span class="EOP Selected SCXW33688343 BCX8" data-ccp-props="{&quot;335559738&quot;:340,&quot;335559739&quot;:140,&quot;335572079&quot;:4,&quot;335572080&quot;:4,&quot;335572081&quot;:13684944,&quot;469789806&quot;:&quot;single&quot;}"> </span></h2>
<p><span data-contrast="none">Ask two people the same business question in different words, and a language model can generate two different SQL statements, and two different numbers. That&#8217;s the failure Metric Views were built to close.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">A Metric View is a Unity Catalog object that defines dimensions, measures, joins, and, in current releases, parameters that let one definition answer differently depending on what&#8217;s calling it: a dashboard, a Genie Agent conversation, or a Genie One skill. Write a readmission_rate_30d calculation once, certify it once, and every surface on the platform reuses that same logic instead of re-deriving it from scratch. 2026 updates added native median and percentile expressions (useful for skewed metrics like length of stay), cluster-by configuration for large fact tables, and wildcard expressions that cut boilerplate when composing layered views.</span></p>
<h2><span class="TextRun SCXW121407507 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">Genie Agent or Custom Predictive Agent: Which </span><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">o</span><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">ne </span><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">d</span><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">o </span><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">y</span><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">ou </span><span class="NormalTextRun AdvancedProofingIssueV2Themed SCXW121407507 BCX8" data-ccp-parastyle="heading 2">a</span><span class="NormalTextRun AdvancedProofingIssueV2Themed SCXW121407507 BCX8" data-ccp-parastyle="heading 2">ctually </span><span class="NormalTextRun AdvancedProofingIssueV2Themed SCXW121407507 BCX8" data-ccp-parastyle="heading 2">n</span><span class="NormalTextRun AdvancedProofingIssueV2Themed SCXW121407507 BCX8" data-ccp-parastyle="heading 2">eed</span><span class="NormalTextRun SCXW121407507 BCX8" data-ccp-parastyle="heading 2">?</span></span><span class="EOP Selected SCXW121407507 BCX8" data-ccp-props="{&quot;335559738&quot;:340,&quot;335559739&quot;:140,&quot;335572079&quot;:4,&quot;335572080&quot;:4,&quot;335572081&quot;:13684944,&quot;469789806&quot;:&quot;single&quot;}"> </span></h2>
<p><span data-contrast="none">A Genie Agent answers questions from data that already exists. &#8220;What&#8217;s our 30-day readmission rate this quarter?&#8221; sits squarely in its lane, even with Agent Mode&#8217;s deeper reasoning. The moment a question turns forward-looking, &#8220;which of my current inpatients is likely to be readmitted, and what should we do about it?&#8221;, you need a custom predictive agent built on the Mosaic AI Agent Framework instead.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">This is a genuinely different tool. You own the model choice, the tool calls, the orchestration, and the evaluation, all inside the same Unity Catalog governance boundary. Genie One can call either one the same way: as an MCP-connected tool, or wrapped inside a skill so a business user never sees the machinery underneath.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<table data-tablestyle="MsoNormalTable" data-tablelook="0" aria-rowcount="6" aria-colcount="3">
<tbody>
<tr aria-rowindex="1">
<td data-celllook="69905"><b><span data-contrast="none">Dimension</span></b><span data-ccp-props="{}"> </span></td>
<td data-celllook="69905"><b><span data-contrast="none">Genie Agent</span></b><span data-ccp-props="{}"> </span></td>
<td data-celllook="69905"><b><span data-contrast="none">Custom Predictive Agent</span></b><span data-ccp-props="{}"> </span></td>
</tr>
<tr aria-rowindex="2">
<td data-celllook="69905"><span data-contrast="none">Best for</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Descriptive, exploratory Q&amp;A over governed tables and Metric Views</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Predictive, multi-step, tool-calling workflows</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="3">
<td data-celllook="69905"><span data-contrast="none">Who owns the logic</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Databricks-managed reasoning and SQL generation</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">You define the reasoning, tools, and orchestration</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="4">
<td data-celllook="69905"><span data-contrast="none">Runs on</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">A SQL Warehouse</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Model Serving endpoints, MLflow models, UC functions</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="5">
<td data-celllook="69905"><span data-contrast="none">Typical output</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">An answer, generated SQL, or a cited Agent Mode report</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">A ranked list, a risk score, or a recommendation</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
<tr aria-rowindex="6">
<td data-celllook="69905"><span data-contrast="none">Reached from Genie One via</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">Direct chat, skills, scheduled tasks</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
<td data-celllook="69905"><span data-contrast="none">An MCP tool connection, or a skill that wraps it</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559740&quot;:260}"> </span></td>
</tr>
</tbody>
</table>
<h2><span class="TextRun SCXW65278092 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">Turning a </span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">r</span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">eadmission </span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">r</span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">isk </span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">s</span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">core </span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">i</span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">nto </span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">a</span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2">ction</span><span class="NormalTextRun SCXW65278092 BCX8" data-ccp-parastyle="heading 2"> in healthcare</span></span><span class="EOP Selected SCXW65278092 BCX8" data-ccp-props="{&quot;335559738&quot;:340,&quot;335559739&quot;:140,&quot;335572079&quot;:4,&quot;335572080&quot;:4,&quot;335572081&quot;:13684944,&quot;469789806&quot;:&quot;single&quot;}"> </span></h2>
<p><span data-contrast="none">Hospital readmissions are one of the most closely watched numbers in healthcare, for good reason. According to CMS, historically about one in five Medicare </span><a href="https://www.cms.gov/medicare/quality/value-based-programs/hospital-readmissions"><span data-contrast="none">patients discharged from a hospital are readmitted within 30 days</span></a><span data-contrast="none">, and the Hospital Readmissions Reduction Program financially penalizes hospitals that exceed their peer benchmark. The gap that matters here isn&#8217;t the prediction. It&#8217;s what happens between a risk score sitting in a dashboard and a care team acting on it before the patient walks out the door.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">A Genie-One-orchestrated version looks like this: </span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<ol>
<li><span data-contrast="auto">A care-coordination lead could define a Genie One </span><b><span data-contrast="none">skill</span></b><span data-contrast="auto"> called </span><span data-contrast="none">weekly_readmission_digest</span><span data-contrast="auto"> that pulls the latest 30-day readmission metrics from a Genie Agent</span></li>
<li><span data-contrast="auto">It cross-references a custom predictive agent&#8217;s high-risk worklist, and formats both into a one-page summary.</span></li>
<li><span data-contrast="auto">A </span><b><span data-contrast="none">scheduled task</span></b><span data-contrast="auto"> then runs that skill every Monday morning and delivers the digest to the unit&#8217;s chat thread and inbox, turning a report someone used to assemble by hand into something that simply shows up, grounded in the same governed data and metric definitions used everywhere else on the platform.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></li>
</ol>
<p><span data-contrast="none">Unifying fragmented clinical data and then acting on it is exactly the kind of work Inferenz does for </span><a href="https://inferenz.ai/industries/healthcare/"><span data-contrast="none">hospital and home-based care operators</span></a><span data-contrast="none"> moving from reactive reporting to real-time, value-based care, using </span><a href="https://inferenz.ai/healthcare-solutions/caregence-predictive-models/"><span data-contrast="none">predictive models built for clinical operations</span></a><span data-contrast="none">.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><i><span data-contrast="none">None of this replaces clinical judgment. Any production system touching protected health information still needs to clear your organization&#8217;s HIPAA and model-governance review before it influences care.</span></i></p>
<h2><span class="TextRun SCXW53474922 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW53474922 BCX8" data-ccp-parastyle="heading 2">Tuning accuracy and cost: The discipline behind reliable answers</span></span><span class="EOP Selected SCXW53474922 BCX8" data-ccp-props="{&quot;335559738&quot;:340,&quot;335559739&quot;:140,&quot;335572079&quot;:4,&quot;335572080&quot;:4,&quot;335572081&quot;:13684944,&quot;469789806&quot;:&quot;single&quot;}"> </span></h2>
<p><span data-contrast="none">Genie One is only as good as the Genie Agents and Metric Views feeding it, so accuracy is a stack-wide habit, not a setting you flip once and forget.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">Start by writing down 20 to 50 representative questions with known-correct answers before touching a single instruction. Prioritize SQL examples and trusted assets over prose instructions, since concrete patterns anchor SQL generation far more reliably than descriptive text. Keep instructions short and free of contradictions, and re-run the benchmark after every schema change or instruction edit. Field reports cite 10 to 40 percent accuracy gains from this loop alone, applied consistently, against an agent nobody ever benchmarks.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<table data-tablestyle="MsoTableGrid" data-tablelook="1184" aria-rowcount="5" aria-colcount="3">
<tbody>
<tr aria-rowindex="1">
<td data-celllook="0"><b><span data-contrast="auto">Result</span></b><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><b><span data-contrast="auto">What It Means</span></b><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><b><span data-contrast="auto">What To Do</span></b><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
</tr>
<tr aria-rowindex="2">
<td data-celllook="0"><span data-contrast="auto">Pass</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Correct answer, correct grain</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Keep as a regression test</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
</tr>
<tr aria-rowindex="3">
<td data-celllook="0"><span data-contrast="auto">Partial</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Right direction, wrong filter or period</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Add a targeted SQL example</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
</tr>
<tr aria-rowindex="4">
<td data-celllook="0"><span data-contrast="auto">Fail, schema</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Can&#8217;t find or join the right tables</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Add column comments, keys, or a Metric View</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
</tr>
<tr aria-rowindex="5">
<td data-celllook="0"><span data-contrast="auto">Fail, ambiguity</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Maps to more than one plausible metric</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
<td data-celllook="0"><span data-contrast="auto">Add a trusted asset to disambiguate</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></td>
</tr>
</tbody>
</table>
<p><span data-contrast="none">Cost follows a similar rhythm. Serverless SQL Warehouses suit the bursty, ad-hoc pattern of conversational analytics better than always-on clusters, and Genie&#8217;s query-level attribution makes it possible to track cost per agent and manage the operating model, not just per warehouse, which is what makes chargeback to a specific business unit realistic. Agent Mode and scheduled tasks both add real compute. Budget for them separately from ad-hoc chat, and audit schedules nobody actually reads.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-16613" src="https://inferenz.ai/wp-content/uploads/2026/08/Rolling-out-an-agentic-coworker-is-a-data-governance-and-change-management-project-1.jpg" alt="Contact our Data and AI Experts" width="1340" height="350" srcset="https://inferenz.ai/wp-content/uploads/2026/08/Rolling-out-an-agentic-coworker-is-a-data-governance-and-change-management-project-1.jpg 1340w, https://inferenz.ai/wp-content/uploads/2026/08/Rolling-out-an-agentic-coworker-is-a-data-governance-and-change-management-project-1-300x78.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/08/Rolling-out-an-agentic-coworker-is-a-data-governance-and-change-management-project-1-1024x267.jpg 1024w, https://inferenz.ai/wp-content/uploads/2026/08/Rolling-out-an-agentic-coworker-is-a-data-governance-and-change-management-project-1-768x201.jpg 768w" sizes="auto, (max-width: 1340px) 100vw, 1340px" /></a></p>
<h2><span class="TextRun SCXW206370826 BCX8" lang="EN-IN" xml:lang="EN-IN" data-contrast="none"><span class="NormalTextRun SCXW206370826 BCX8" data-ccp-parastyle="heading 2">The </span><span class="NormalTextRun SCXW206370826 BCX8" data-ccp-parastyle="heading 2">b</span><span class="NormalTextRun SCXW206370826 BCX8" data-ccp-parastyle="heading 2">ottom </span><span class="NormalTextRun SCXW206370826 BCX8" data-ccp-parastyle="heading 2">l</span><span class="NormalTextRun SCXW206370826 BCX8" data-ccp-parastyle="heading 2">ine</span></span><span class="EOP Selected SCXW206370826 BCX8" data-ccp-props="{&quot;335559738&quot;:340,&quot;335559739&quot;:140,&quot;335572079&quot;:4,&quot;335572080&quot;:4,&quot;335572081&quot;:13684944,&quot;469789806&quot;:&quot;single&quot;}"> </span></h2>
<p><span data-contrast="none">Genie One gives every business team one coworker to talk to, instead of five dashboards and an analyst&#8217;s calendar. But the chat interface is the easy part. What makes it trustworthy is everything underneath: Genie Agents that turn governed Unity Catalog data into plain language, Metric Views that keep every surface computing the same number, and custom predictive agents that pick up exactly where descriptive analytics runs out of road.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<p><span data-contrast="none">Treat the whole stack the way you would any production system: benchmark it, tune it on a short loop, and keep watching it after launch. The organizations already ahead here aren&#8217;t the ones with the flashiest chat interface. They are the ones who did the </span><a href="https://inferenz.ai/services/data-and-cloud-modernization/"><span data-contrast="none">unglamorous data foundation work</span></a><span data-contrast="none"> first.</span><span data-ccp-props="{&quot;201341983&quot;:0,&quot;335559739&quot;:160,&quot;335559740&quot;:300}"> </span></p>
<h2 aria-level="2"><span data-contrast="none">Frequently Asked Questions</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p>The post <a href="https://inferenz.ai/blogs/databricks-genie-one-inside-the-agentic-coworker-turning-business-data-into-action/">Databricks Genie One: Inside the Agentic Coworker Turning Business Data into Action</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Parquet v2 in Azure Databricks: What Changed and Why It Matters</title>
		<link>https://inferenz.ai/blogs/parquet-v2-in-azure-databricks-what-changed-and-why-it-matters/</link>
		
		<dc:creator><![CDATA[spectrics]]></dc:creator>
		<pubDate>Fri, 17 Jul 2026 08:07:50 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://inferenz.ai/blogs//</guid>

					<description><![CDATA[<p>Parquet file format v2 is now generally available for Delta Lake and Apache Iceberg tables in Azure Databricks Runtime 18.1 and above. </p>
<p>The post <a href="https://inferenz.ai/blogs/parquet-v2-in-azure-databricks-what-changed-and-why-it-matters/">Parquet v2 in Azure Databricks: What Changed and Why It Matters</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h2><b><span data-contrast="auto">Summary</span></b><span data-contrast="auto"> </span><span data-ccp-props="{}"> </span></h2>
<p><strong>Parquet file format</strong> v2 is now generally available for Delta Lake and Apache Iceberg tables in Azure Databricks Runtime 18.1 and above. It swaps in smarter encodings, richer page-level metadata, and INT64 timestamps to shrink file sizes and speed up querieswith zero changes to your existing SQL. Turn it on with a single table property and use REORG TABLE when you want your historical data rewritten too.</p>
<h2>Introduction</h2>
<p>Global data volumes are on pace to cross roughly 230–240 zettabytes by 2026, according to Statista estimates, and every terabyte of that sits somewhere, on someone&#8217;s storage bill. Most of it, if you&#8217;re running an Azure Databricks lakehouse, sits in Parquet file format files. So when the format underneath your Delta Lake or Apache Iceberg tables gets a meaningful upgrade, it&#8217;s worth fifteen minutes of your attention.</p>
<p>That upgrade is Parquet v2, and unlike a major platform migration, it doesn&#8217;t ask you to touch a single line of application code or rewrite a query. It changes how data is physically packed inside the files themselves which, in practice, means smaller files, better compression, and faster reads for the same data you already have.</p>
<p><strong>A quick note on terms, if you&#8217;re newer to the stack: </strong>Apache Parquet is the columnar file format that stores data by column rather than by row, letting a query engine like Spark read only the columns a query actually needs. Delta Lake is the transactional layer Databricks builds on top of Parquet it adds ACID guarantees, schema enforcement, and time travel. Apache Iceberg is a similar open table format, increasingly used alongside or instead of Delta Lake in mixed-engine environments. Parquet v2 sits one level below both of them, at the file format itself, which is exactly why it works across both table types without any application-level rework.</p>
<p>In this <strong>Parquet v2 Azure Databricks</strong> guide, we&#8217;ll cover what Parquet v2 actually changes compared to Parquet v1, how to turn it on, what to check before you do, and where it fits in a broader Databricks storage optimization strategy.</p>
<h2>What Is Parquet v2?</h2>
<p><strong>Parquet file format</strong>has been the default columnar storage format for big data platforms for well over a decade, and for good reason, it stores data efficiently and let’s query engines skip columns a query doesn&#8217;t touch. In wide tables, that column-pruning alone can cut I/O dramatically; reading two columns out of a hundred means the engine never has to touch the other ninety-eight.</p>
<p>But the original <strong>Parquet file format</strong>spec (v1) was designed for a different era of data volumes, and three limitations became increasingly visible as workloads scaled:</p>
<ul>
<li>Integer and string compression left performance on the table.</li>
<li>Query engines had limited page-level metadata to use for skipping unnecessary data.</li>
<li>Timestamps relied on the older INT96 format, which compressed and filtered poorly.</li>
</ul>
<p>Parquet v2 addresses all three with more efficient encodings, richer page metadata, and a modern INT64 timestamp format. According to Microsoft&#8217;s Azure Databricks documentation and Databricks&#8217; own platform release notes, <strong>Parquet v2 is generally available for Delta Lake and Apache Iceberg tables starting in Databricks Runtime 18.1</strong>, with support for converting existing data added in Runtime 18.2.</p>
<h3>Parquet v1 vs. Parquet v2, at a Glance</h3>
<table width="624">
<tbody>
<tr>
<td width="208"><strong>Capability</strong></td>
<td width="208"><strong>Parquet v1</strong></td>
<td width="208"><strong>Parquet v2</strong></td>
</tr>
<tr>
<td width="208">Timestamp storage</td>
<td width="208">INT96</td>
<td width="208">INT64 (better compression, more accurate stats)</td>
</tr>
<tr>
<td width="208">Integer / string encoding</td>
<td width="208">Standard RLE / dictionary</td>
<td width="208">Adds DELTA_BINARY_PACKED and DELTA_LENGTH_BYTE_ARRAY for tighter packing</td>
</tr>
<tr>
<td width="208">Page metadata</td>
<td width="208">Basic headers</td>
<td width="208">Richer v2 data page headers with per-page stats</td>
</tr>
<tr>
<td width="208">Predicate pushdown</td>
<td width="208">Limited page-level skipping</td>
<td width="208">Finer-grained data skipping at the page level</td>
</tr>
<tr>
<td width="208">Enable via</td>
<td width="208">Default</td>
<td width="208">delta.parquet.format.version / iceberg.parquet.format.version = 2.12.0</td>
</tr>
<tr>
<td width="208">Reader compatibility</td>
<td width="208">Universal</td>
<td width="208">Broad, but verify external / third-party readers</td>
</tr>
</tbody>
</table>
<h2>What&#8217;s New in Parquet v2?</h2>
<h3>1. Better Compression for Integers and Strings</h3>
<p>The headline change is how numeric and string values get stored. Parquet v2 introduces more efficient encoding techniques, including delta-based packing for integers and byte arrays, per the Apache Parquet specification, which typically produce:</p>
<ul>
<li>Smaller Parquet files</li>
<li>Better compression ratios</li>
<li>Faster decoding during query execution</li>
</ul>
<p>Independent benchmarking from the DuckDB engineering team gives a useful sense of scale here: enabling Parquet v2&#8217;s newer encodings produced files roughly 30% smaller with 15% faster writes under Snappy compression, and about 11% smaller with 24% faster writes under zstd, across their test datasets. On highly sequential data, think auto-incrementing IDs or evenly spaced timestamps, the gains were far more dramatic, with some columns shrinking by over 90%.</p>
<p><strong>One honest caveat </strong>worth flagging: delta encoding isn&#8217;t a universal win. On columns with moderate entropydata that repeats in patterns but isn&#8217;t cleanly sequential, delta encoding can occasionally produce larger files than v1, because it turns repeating values into effectively random deltas that compress worse. It&#8217;s a good reason to test against a representative sample of your own tables rather than assuming uniform gains.</p>
<h3>2. Improved Data Page Headers</h3>
<p>Parquet files are divided into pages, and in Parquet v2, each page carries richer metadata, allowing Databricks to determine whether a page contains relevant data before it&#8217;s ever read. That directly improves:</p>
<ul>
<li>Predicate pushdown</li>
<li>Data skipping</li>
<li>Query performance on filtered datasets</li>
</ul>
<p>If your query filters sales data for a single month, Databricks can skip the pages that don&#8217;t contain data from that period, cutting the volume of data scanned, and the compute cost that comes with it.</p>
<h3>3. INT64 Timestamps Instead of INT96</h3>
<p>Older Parquet files stored timestamps in the INT96 format. Parquet v2 replaces it with the more efficient INT64 timestamp format, which brings:</p>
<ul>
<li>Better compression</li>
<li>More accurate statistics</li>
<li>Faster filtering on timestamp columns</li>
<li>Better compatibility with modern analytics engines</li>
</ul>
<p>Since most enterprise data warehouse tables lean heavily on timestamp columns event logs, transaction records, IoT telemetry this single change tends to have an outsized, noticeable impact on real-world query performance.</p>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignleft size-full wp-image-15996" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2026/07/Not-sure-if-your-tables-are-ready-for-Parquet-v2.jpg" alt="" width="870" height="235" srcset="https://inferenz.ai/wp-content/uploads/2026/07/Not-sure-if-your-tables-are-ready-for-Parquet-v2.jpg 870w, https://inferenz.ai/wp-content/uploads/2026/07/Not-sure-if-your-tables-are-ready-for-Parquet-v2-300x81.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/07/Not-sure-if-your-tables-are-ready-for-Parquet-v2-768x207.jpg 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></a></p>
<h2>How to Enable Parquet v2</h2>
<p>If you&#8217;re using <a href="https://inferenz.ai/blogs/databricks-unity-catalog-building-a-unified-data-governance-layer-in-modern-data-platforms/">Databricks Unity Catalog</a> managed tables, Azure Databricks may automatically upgrade compatible tables to Parquet v2 for you. To enable it manually, set the table property directly.</p>
<p><strong>Existing Delta table</strong><br />
<code><br />
ALTER TABLE table_name<br />
SET TBLPROPERTIES (<br />
'delta.parquet.format.version' = '2.12.0'<br />
);<br />
</code><br />
<strong>Existing Iceberg table</strong><br />
<code><br />
ALTER TABLE table_name<br />
SET TBLPROPERTIES (<br />
'iceberg.parquet.format.version' = '2.12.0'<br />
);<br />
</code><br />
<strong>New Delta table</strong><br />
<code><br />
CREATE TABLE table_name (...)<br />
TBLPROPERTIES (<br />
'delta.parquet.format.version' = '2.12.0'<br />
);<br />
</code><br />
<strong>New Iceberg table</strong><br />
<code><br />
CREATE TABLE table_name (...)<br />
USING iceberg<br />
TBLPROPERTIES (<br />
'iceberg.parquet.format.version' = '2.12.0'<br />
);<br />
</code></p>
<h2>Existing Data Isn&#8217;t Automatically Converted</h2>
<p>This is the detail most teams miss on their first pass.</p>
<p>Changing the table property only affects new data written after the change. Your existing Parquet files stay in their original format, which means a single table can temporarily hold a mix of Parquet v1 and Parquet v2 files side by side.</p>
<p>If you want your historical data converted too, Databricks Runtime 18.2 and above provides the REORG TABLE command:<br />
<code><br />
REORG TABLE table_name<br />
APPLY (<br />
SET PARQUET (FORMAT_VERSION = '2.12.0')<br />
);<br />
</code></p>
<p>This rewrites every existing file using Parquet v2, so the entire table benefits from the format change.</p>
<h2>Can You Roll Back?</h2>
<p>Yes. If you hit a compatibility issue downstream, you can convert the table back to Parquet v1 just as easily:</p>
<p><code><br />
REORG TABLE table_name<br />
APPLY (<br />
SET PARQUET (FORMAT_VERSION = '1.0.0')<br />
);<br />
</code></p>
<p>This rewrites the data files again and restores the table to the older format a genuinely low-risk way to test Parquet v2 in a non-production environment before committing.</p>
<h2>Things to Check Before Enabling Parquet v2</h2>
<p>Parquet v2 delivers real gains, but it&#8217;s worth verifying compatibility if your data is read outside Databricks. Before flipping the switch, check for:</p>
<ul>
<li><strong>External Apache Iceberg readers </strong>that may not yet support Parquet v2. DuckDB&#8217;s own engineering team, for instance, has noted that several mainstream query engines still default to writing (and in some cases reading) older Parquet encodings for exactly this reason.</li>
<li><strong>Delta Sharing or other external sharing methods</strong> confirm recipient tools can actually read Parquet v2 files before you share.</li>
<li><strong>Materialized views and streaming tables</strong>, which aren&#8217;t upgraded automatically and need to be enabled manually.</li>
</ul>
<p>If your tables are read exclusively from within Databricks, compatibility generally isn&#8217;t a concern at all.</p>
<h2>Should You Use Parquet v2?</h2>
<p>For most Azure Databricks workloads, yes. Parquet v2 offers real advantages without requiring any application-side changes:</p>
<ul>
<li>Reduced storage usage</li>
<li>Better compression</li>
<li>Faster query execution</li>
<li>Improved predicate pushdown</li>
<li>Better timestamp handling</li>
</ul>
<p>If your data is read exclusively within Azure Databricks, enabling Parquet v2 is a low-effort, high-leverage way to improve performance. If external tools, BI platforms, or third-party engines also touch your data, verify compatibility first.</p>
<p>A practical rollout sequence looks like this:</p>
<ol>
<li>Let <strong>Databricks Unity Catalog</strong> managed tables upgrade automatically where applicable.</li>
<li>Enable Parquet v2 on Databricks-only tables first.</li>
<li>Run REORG TABLE if you want existing data converted.</li>
<li>Test external readers before enabling Parquet v2 on shared datasets.</li>
</ol>
<h2>Final Thoughts</h2>
<p>Parquet v2 is the kind of improvement that works quietly in the background. You keep writing the same SQL, running the same pipelines, using the same BI tools but your data becomes more storage-efficient, and your queries, more often than not, come back a little faster.</p>
<p>For enterprises running Azure Databricks at scale, this is a low-effort upgrade with a real payoff, provided you verify compatibility for anyone reading your data outside the platform. As a Databricks consulting partner, Inferenz has helped enterprise data teams evaluate exactly this kind of platform-level change as part of broader Delta Lake and lakehouse cost-optimization engagements. Through its <a href="https://inferenz.ai/services/data-engineering-and-integration/">data engineering and integration services</a>, organizations have optimized storage architectures, improved data performance, and reduced cloud infrastructure costs, where a one-line table property, applied correctly across the right tables, adds up to a measurable line item on the cloud bill.</p>
<p><a href=" https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignleft wp-image-15995 size-full" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2026/07/Your-Databricks-storage-costs-wont-optimize-themselves.jpg" alt="" width="870" height="235" srcset="https://inferenz.ai/wp-content/uploads/2026/07/Your-Databricks-storage-costs-wont-optimize-themselves.jpg 870w, https://inferenz.ai/wp-content/uploads/2026/07/Your-Databricks-storage-costs-wont-optimize-themselves-300x81.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/07/Your-Databricks-storage-costs-wont-optimize-themselves-768x207.jpg 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></a></p>
<h2>Frequently Asked Questions</h2>
<p>The post <a href="https://inferenz.ai/blogs/parquet-v2-in-azure-databricks-what-changed-and-why-it-matters/">Parquet v2 in Azure Databricks: What Changed and Why It Matters</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Databricks Data + AI Summit 2026: The Lakehouse Just Became Something Bigger</title>
		<link>https://inferenz.ai/blogs/databricks-data-ai-summit-2026-the-lakehouse-just-became-something-bigger/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Wed, 08 Jul 2026 05:20:42 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://inferenz.ai/blogs//</guid>

					<description><![CDATA[<p>Databricks Data + AI Summit 2026 was not a feature release, but a declaration. The Lakehouse is no longer just where enterprises store and query data.</p>
<p>The post <a href="https://inferenz.ai/blogs/databricks-data-ai-summit-2026-the-lakehouse-just-became-something-bigger/">Databricks Data + AI Summit 2026: The Lakehouse Just Became Something Bigger</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h2><b><span data-contrast="auto">Summary</span></b><span data-ccp-props="{}"> </span></h2>
<p><span data-contrast="auto">Databricks Data + AI Summit 2026 was not a feature release, but a declaration. The Lakehouse is no longer just where enterprises store and query data. It is where agents do the job for your business. Here is what changed, what it means, and why it matters now.</span><span data-ccp-props="{}"> </span></p>
<h2 aria-level="2"><span data-contrast="none">Introduction</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">Every year, the tech industry produces a hundred summits that announce things.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">Databricks Data and AI Summit 2026 (DAIS) was different. </span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">What Databricks put on the table in San Francisco this June was an architectural argument about where enterprise AI is actually headed, and it landed with the kind of coherence that makes you reconsider how you have been thinking about your data stack.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">The theme, if you had to name it, was this: the </span><b><i><span data-contrast="auto">Lakehouse is now the control plane for the agentic enterprise</span></i></b><span data-contrast="auto">. Not just a place to store data. The place where agents govern, reason, act, and get held accountable for what they do.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">For Inferenz, a Databricks partner </span><b><span data-contrast="auto">building agentic AI solutions in healthcare</span></b><span data-contrast="auto"> and enterprise, several of these announcements land directly in the infrastructure we build on and deploy for clients. </span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">Here is our read on what mattered most and what you should actually do about it.</span><span data-ccp-props="{}"> </span></p>
<h2 aria-level="2"><span data-contrast="none">The context problem is finally being taken seriously</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">If you have ever deployed an AI model and watched it produce a confidently wrong answer, you already know the core problem DAIS 2026 addressed.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">It is not model quality. It is context.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">As Ali Ghodsi, founder and CEO of Databricks put it during the keynote: &#8220;Most enterprise AI today is just guessing with false confidence. If you&#8217;re a CFO and AI can&#8217;t tell you why margins changed, that&#8217;s not an AI problem. That&#8217;s a context problem.&#8221;</span><span data-ccp-props="{}"> </span></p>
<p><b><span data-contrast="auto">Genie Ontology</span></b><span data-contrast="auto"> is Databricks&#8217;s answer to that. It is a live context layer that continuously reads your data, documents, queries, and applications to build a machine-readable map of what your business actually means by its own terms. </span><span data-ccp-props="{}"> </span></p>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="7" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="1" data-aria-level="1"><span data-contrast="auto">What does &#8220;active user&#8221; mean in your system? </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="7" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="2" data-aria-level="1"><span data-contrast="auto">What is your definition of &#8220;churn&#8221;? </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="7" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="3" data-aria-level="1"><span data-contrast="auto">When did your ARR calculation change, and why?</span><span data-ccp-props="{}"> </span></li>
</ul>
<p><span data-contrast="auto">This is not a static data dictionary someone fills in once and forgets. Genie Ontology updates continuously, weighs sources by authority (similar to how PageRank works), and feeds that knowledge directly into Unity Catalog&#8217;s semantic layer. </span><span data-ccp-props="{}"> </span></p>
<p><i><span data-contrast="auto">The downstream effect:</span></i><span data-contrast="auto"> every agent, every dashboard, and every AI-generated report pulls from one shared, authoritative understanding of your business rather than each making its own guesses.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">The company with the best context layer will have a larger AI advantage than the company with the most data. That sentence from the Bain team covering the summit deserves to sit with you for a moment.</span><span data-ccp-props="{}"> </span></p>
<h2 aria-level="3"><span data-contrast="none">Genie One: an AI coworker that actually knows your business</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><b><span data-contrast="auto">Genie One</span></b><span data-contrast="auto"> is now generally available, and it is a significant step past what most enterprise AI assistants can actually do.</span><span data-ccp-props="{}"> </span></p>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="2" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="15" data-aria-level="1"><span data-contrast="auto">It connects to over 50 applications, including Gmail, Slack, Teams, Jira, and Confluence. </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="2" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="16" data-aria-level="1"><span data-contrast="auto">It can answer questions grounded in your actual governed lakehouse data, draft documents, schedule tasks, monitor changes, and explain why something happened. </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="2" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="17" data-aria-level="1"><span data-contrast="auto">On a benchmark of 28 real-world enterprise data questions, Genie answered 84.5% correctly on the first attempt. The best general-purpose coding agent on the same test scored 52.4%.</span><span data-ccp-props="{}"> </span></li>
</ul>
<p><span data-contrast="auto">The difference is the ontology layer underneath. Genie is not searching documents. It is reasoning against a live, governed representation of your business. That is what separates a useful answer from a plausible one.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">For enterprise teams evaluating where to start with agentic AI, Genie One is the fastest path to ROI for non-technical business users. No seat-based pricing, either. Each user gets 150 DBUs of free LLM usage per month, with pay-as-you-go beyond that.</span><span data-ccp-props="{}"> </span></p>
<h2 aria-level="2"><span data-contrast="none">LTAP: Forty years of infrastructure debt, addressed</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">Here is a problem most enterprises have accepted as permanent: your transactional systems and your analytical systems have always been two separate things. Separate databases, separate formats, ETL pipelines running between them, two slightly different copies of the same data that never quite agreed.</span><span data-ccp-props="{}"> </span></p>
<p><b><span data-contrast="auto">LTAP (Lake Transactional/Analytical Processing)</span></b><span data-contrast="auto"> changes that. </span><span data-ccp-props="{}"> </span></p>
<p><i><span data-contrast="auto">The mechanics:</span></i><span data-contrast="auto"> </span><b><span data-contrast="auto">Lakebase</span></b><span data-contrast="auto">, Databricks&#8217;s serverless PostgreSQL database (now at 12 million launches per day), stores transactional data directly in Unity Catalog using Delta and Iceberg formats. </span><span data-ccp-props="{}"> </span></p>
<p><span data-ccp-props="{}">No ETL. No sync. Hidden copies disappear. Every analytical engine reads the same governed file. </span></p>
<p><span data-contrast="auto">For AI agents, this is foundational. An agent that needs to read a customer&#8217;s live order history and then run six months of purchasing analysis currently must query two systems and reconcile two copies of data. With LTAP, there is one copy, one governance layer, and one point of truth.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">New Lakebase capabilities at the summit: </span><span data-ccp-props="{}"> </span></p>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="2" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="18" data-aria-level="1"><span data-contrast="auto">cross-cloud disaster recovery</span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="2" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="19" data-aria-level="1"><span data-contrast="auto">git-style database branching (spin up a full-fidelity clone of production in sub-seconds for safe testing), and </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="2" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="20" data-aria-level="1"><span data-contrast="auto">Lakebase Search, which brings hybrid vector and full-text retrieval natively into PostGRES.</span><span data-ccp-props="{}"> </span></li>
</ul>
<p><b><span data-contrast="auto">Lakehouse//RT</span></b><span data-contrast="auto">, powered by a new engine called Reyden, rounds this out with sub-100ms query latency at 12,000 queries per second directly on Delta and Iceberg tables. PointClickCare&#8217;s benchmarks showed it running more than a third faster than their prior warehouse, on their own healthcare dataset, without a separate serving system.</span><span data-ccp-props="{}"> </span></p>
<h2 aria-level="2"><span data-contrast="none">Agent Bricks: The platform that does the other 99%</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">Building an AI agent is not hard anymore. The hard part is everything else.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">Memory across sessions. Security when agents execute code. Cost management when agents run at scale. Evaluation. Monitoring. Governance of what they can access. That is the 99% of engineering work that does not show up in demos but determines whether your deployment works in production.</span><span data-ccp-props="{}"> </span></p>
<p><b><span data-contrast="auto">Agent Bricks</span></b><span data-contrast="auto"> is now a full-stack platform for exactly that. Over 100,000 agents have been built on it. AstraZeneca, 7-Eleven, Fox, and Block all run production agents on Agent Bricks. The 2026 expansion added:</span><span data-ccp-props="{}"> </span></p>
<ul>
<li aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Symbol&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;&quot;,&quot;469777815&quot;:&quot;multilevel&quot;}" data-aria-posinset="1" data-aria-level="1"><span data-contrast="auto">Managed agent memory powered by Lakebase, persistent across sessions</span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Symbol&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;&quot;,&quot;469777815&quot;:&quot;multilevel&quot;}" data-aria-posinset="2" data-aria-level="1"><span data-contrast="auto">MCP-connected retrieval from Unity Catalog and external tools like GitHub, Jira, and Google Drive</span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Symbol&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;&quot;,&quot;469777815&quot;:&quot;multilevel&quot;}" data-aria-posinset="3" data-aria-level="1"><span data-contrast="auto">Secure sandboxed compute environments for code execution</span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Symbol&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;&quot;,&quot;469777815&quot;:&quot;multilevel&quot;}" data-aria-posinset="4" data-aria-level="1"><span data-contrast="auto">Multi-model support: OpenAI, Anthropic, Gemini, Qwen, Grok, all governed under Unity Catalog</span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="" data-font="Symbol" data-listid="1" data-list-defn-props="{&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Symbol&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;&quot;,&quot;469777815&quot;:&quot;multilevel&quot;}" data-aria-posinset="5" data-aria-level="1"><b><span data-contrast="auto">Omnigent</span></b><span data-contrast="auto">, a meta-orchestration layer for managing agents across different frameworks, models, and tools when your stack is not monolithic</span><span data-ccp-props="{}"> </span></li>
</ul>
<p><span data-contrast="auto">For teams building with Claude Code SDK, LangGraph, CrewAI, or OpenAI Agent SDKs, Omnigent is the layer that lets these coexist under one governance model instead of sprawling across disconnected stacks.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">Databricks also moved five products into the free tier: Genie Code, Serverless GPUs, Lakebase, Agent Bricks, and Lakeflow Designer. You can now prototype an entire agentic application from data pipeline to agent logic to served endpoint without spending anything.</span><span data-ccp-props="{}"> </span></p>
<h2 aria-level="2"><span data-contrast="none">Unity AI Gateway: Governance that happens at runtime</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">This is the announcement that regulated industries have been waiting for.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">Traditional AI governance asked: who can access which data, which model is approved? That works for humans. It breaks down for agents that act autonomously, spawn subagents, call external tools, and generate outputs at volume.</span><span data-ccp-props="{}"> </span></p>
<p><b><span data-contrast="auto">Unity AI Gateway</span></b><span data-contrast="auto"> governs what agents actually do at the moment they do it. </span><span data-ccp-props="{}"> </span></p>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="8" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="15" data-aria-level="1"><span data-contrast="auto">Hard spend caps. </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="8" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="16" data-aria-level="1"><span data-contrast="auto">Real-time PII detection. </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="8" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="17" data-aria-level="1"><span data-contrast="auto">Prompt injection prevention. </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="8" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="18" data-aria-level="1"><span data-contrast="auto">Full trace capture of every tool call, MCP interaction, and subagent action. </span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="8" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="19" data-aria-level="1"><span data-contrast="auto">Security policies written in SQL that respond to agent behavior in context, not just static rules applied at the edge.</span><span data-ccp-props="{}"> </span></li>
</ul>
<p><span data-contrast="auto">At Inferenz, our work in healthcare AI has always required governance to be a first-class concern. What Unity AI Gateway represents is exactly the infrastructure required to move agentic AI from pilot deployments into production clinical environments. </span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">We cover this in more depth through our </span><a href="https://inferenz.ai/services/generative-and-agentic-ai/"><span data-contrast="none">Generative and Agentic AI services</span></a><span data-contrast="auto"> and the governance architecture that underpins our </span><a href="https://inferenz.ai/healthcare-solutions/caregence-platform/"><span data-contrast="none">Caregence platform</span></a><span data-contrast="auto">.</span><span data-ccp-props="{}"> </span></p>
<p><a href="https://inferenz.ai/healthcare-solutions/caregence-agents/"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-15772" src="https://inferenz.ai/wp-content/uploads/2026/07/Is-Your-AI-Infrastructure-Ready-for-Agent-Scale-Governance.jpg" alt="Is-Your-AI-Infrastructure-Ready-for-Agent-Scale-Governance" width="870" height="235" srcset="https://inferenz.ai/wp-content/uploads/2026/07/Is-Your-AI-Infrastructure-Ready-for-Agent-Scale-Governance.jpg 870w, https://inferenz.ai/wp-content/uploads/2026/07/Is-Your-AI-Infrastructure-Ready-for-Agent-Scale-Governance-300x81.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/07/Is-Your-AI-Infrastructure-Ready-for-Agent-Scale-Governance-768x207.jpg 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></a></p>
<h2 aria-level="2"><span data-contrast="none">OpenSharing: Open standards win again!</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">Databricks launched Delta Sharing in 2021 to solve cross-organizational data sharing without copying files. It became the most widely adopted open data-sharing protocol in the industry.</span><span data-ccp-props="{}"> </span></p>
<p><b><span data-contrast="auto">OpenSharing</span></b><span data-contrast="auto"> extends that logic to the full AI stack. Data, models, agent skills, and Genie Agents can now be shared across organizations and clouds via a single Linux Foundation-hosted open protocol.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">The practical enterprise use case is Genie Agent Sharing: share a governed AI interface with a partner or customer, giving them curated access to your data and reasoning capabilities without exposing your underlying logic, proprietary calculations, or source tables. You control what they can ask, how much data they can export, and how many requests they can make.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">SecureConnect removes the networking headache: cross-cloud storage connections without per-recipient firewall configuration.</span><span data-ccp-props="{}"> </span></p>
<h2 aria-level="2"><span data-contrast="none">What this means for Inferenz clients</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto"><a href="https://inferenz.ai/news/inferenz-partners-with-databricksto-drive-data-ai-and-generative-ai/">Inferenz is a Databricks partner.</a> We build on this platform. Several of the Databricks summit announcements directly expand what we can deliver:</span><span data-ccp-props="{}"> </span></p>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="3" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="15" data-aria-level="1"><b><span data-contrast="auto">Genie Ontology</span></b><span data-contrast="auto"> strengthens the semantic layer that our healthcare clients need for AI to reason correctly about clinical terms, payer rules, and care metrics without every agent reinventing the definition.</span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="3" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="16" data-aria-level="1"><b><span data-contrast="auto">Lakebase and LTAP</span></b><span data-contrast="auto"> close the gap between transactional care data and the analytical models that power Caregence predictive risk intelligence. Patient records that update in real time can now feed directly into risk models without ETL delays.</span><span data-ccp-props="{}"> </span></li>
</ul>
<ul>
<li aria-setsize="-1" data-leveltext="-" data-font="Aptos" data-listid="3" data-list-defn-props="{&quot;335551671&quot;:15,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Aptos&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;-&quot;,&quot;469777815&quot;:&quot;hybridMultilevel&quot;}" data-aria-posinset="17" data-aria-level="1"><b><span data-contrast="auto">Agent Bricks governance and Unity AI Gateway</span></b><span data-contrast="auto"> provide the runtime controls our healthcare deployments require. HIPAA-compliant agentic AI is not just a compliance checkbox. It is an architecture. These capabilities make that architecture standard rather than custom-built for every engagement.</span><span data-ccp-props="{}"> </span></li>
</ul>
<p><span data-contrast="auto">For enterprise clients working on </span><a href="https://inferenz.ai/services/data-and-cloud-modernization/"><span data-contrast="none">data and cloud modernization</span></a><span data-contrast="auto"> or evaluating where agentic AI fits in their stack, the LTAP architecture eliminates an entire tier of infrastructure that was previously unavoidable. One governed copy of data, one permission model, one source of truth for both operational and analytical AI.</span></p>
<h2 aria-level="2"><span data-contrast="none">Five things worth acting on now</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">Most enterprises left the summit with a list of things to watch. These five are worth starting this quarter.</span><span data-ccp-props="{}"> </span></p>
<h3><span data-contrast="none">Define your semantic layer before your agents do it for you. </span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h3>
<p><span data-contrast="auto">Genie Ontology is only as good as what Unity Catalog already knows. If your organization has never agreed on what &#8220;revenue&#8221; or &#8220;active user&#8221; officially means, that conversation is now blocking your AI roadmap.</span><span data-ccp-props="{&quot;335559685&quot;:720}"> </span></p>
<h3><span data-contrast="none">Consolidate your database tier. </span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h3>
<p><span data-contrast="auto">Running a separate operational database alongside Databricks? Lakebase and LTAP give you a clear path to one governed system. The git-style branching alone makes the evaluation worth an afternoon.</span><span data-ccp-props="{&quot;335559685&quot;:720}"> </span></p>
<h3><span data-contrast="none">Audit your agent governance. </span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h3>
<p><span data-contrast="auto">Most AI pilots have no runtime enforcement. If your governance stops at the data catalog, it is not governance. Unity AI Gateway fixes that, but only if you implement it.</span><span data-ccp-props="{&quot;335559685&quot;:720}"> </span></p>
<h3><span data-contrast="none">Prototype on Agent Bricks before building custom. </span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h3>
<p><span data-contrast="auto">Lakebase, Agent Bricks, and Serverless GPUs are all free tier now. There is no budget justification for building a custom agentic stack before you have tested what is already there.</span><span data-ccp-props="{&quot;335559685&quot;:720}"> </span></p>
<h3><span data-contrast="none">Treat context as a strategic asset. </span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h3>
<p><span data-contrast="auto">The next AI advantage will not come from model selection. It will come from the organization whose agents have the clearest, most authoritative understanding of what the business means. That is a semantic architecture decision, not a procurement one.</span></p>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-15773" src="https://inferenz.ai/wp-content/uploads/2026/07/What-Would-Your-Data-Stack-Look-Like-If-It-Was-Built-for-Agents.jpg" alt="What-Would-Your-Data-Stack-Look-Like-If-It-Was-Built-for-Agents" width="870" height="235" srcset="https://inferenz.ai/wp-content/uploads/2026/07/What-Would-Your-Data-Stack-Look-Like-If-It-Was-Built-for-Agents.jpg 870w, https://inferenz.ai/wp-content/uploads/2026/07/What-Would-Your-Data-Stack-Look-Like-If-It-Was-Built-for-Agents-300x81.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/07/What-Would-Your-Data-Stack-Look-Like-If-It-Was-Built-for-Agents-768x207.jpg 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></a></p>
<h2 aria-level="2"><span data-contrast="none">Final thought</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p><span data-contrast="auto">The debate in enterprise AI used to be about which model to choose. DAIS 2026 made clear that this was always the wrong question.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">The model is not the constraint. The architecture around it is.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">Context, governance, live data, and runtime control are the infrastructure that determines whether your AI delivers or stalls. Databricks built a year&#8217;s worth of announcements around exactly those four things.</span><span data-ccp-props="{}"> </span></p>
<p><span data-contrast="auto">For enterprises that have been waiting for the infrastructure to catch up to the ambition, it just did.</span></p>
<h2 aria-level="2"><span data-contrast="none">Frequently Asked Questions</span><span data-ccp-props="{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:160,&quot;335559739&quot;:80}"> </span></h2>
<p>The post <a href="https://inferenz.ai/blogs/databricks-data-ai-summit-2026-the-lakehouse-just-became-something-bigger/">Databricks Data + AI Summit 2026: The Lakehouse Just Became Something Bigger</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>The Home Health Data Visibility Problem and the AI Agents that you Need</title>
		<link>https://inferenz.ai/blogs/the-home-health-data-visibility-problem-and-the-ai-agents-that-you-need/</link>
		
		<dc:creator><![CDATA[spectrics]]></dc:creator>
		<pubDate>Wed, 27 May 2026 11:30:11 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Data & Cloud Migration]]></category>
		<category><![CDATA[Healthcare]]></category>
		<category><![CDATA[Home-Based Care]]></category>
		<guid isPermaLink="false">https://inferenz.ai/blogs//</guid>

					<description><![CDATA[<p>Home health generates more clinical data per patient than almost any other care setting, yet readmissions that remain preventable, keep happening and caregiver turnover sits at 75%. The problem has never been data shortage but data visibility to see patient data as a whole.</p>
<p>The post <a href="https://inferenz.ai/blogs/the-home-health-data-visibility-problem-and-the-ai-agents-that-you-need/">The Home Health Data Visibility Problem and the AI Agents that you Need</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h2><strong>Summary</strong></h2>
<p><em>Home health generates more clinical data per patient than almost any other care setting, yet readmissions that remain preventable, keep happening and caregiver turnover sits at 75%. The problem has never been data shortage but data visibility to see patient data as a whole. </em></p>
<p><em>The <a href="https://inferenz.ai/healthcare-solutions/mpi-and-patient-360/">Master Patient Index &amp; Patient 360 solution </a></em><em>from Inferenz fixes this by resolving fragmented patient identities into one governed record, then builds a chronological Patient 360 timeline on top of it.</em></p>
<h2>The Real Problem Is Not Data. It Is the Architecture.</h2>
<p>I have sat across from enough home health executives to know that &#8220;we don&#8217;t have the data&#8221; is rarely the actual complaint. What they say, when you press them, is closer to: &#8220;We have all this data, and I still can&#8217;t tell you which patients are trending toward hospitalization this week.&#8221;</p>
<p>That is a data architecture problem, not an absence of some clinical system or tool.</p>
<p>The average home health patient generates events across multiple, separate platforms in a single week.</p>
<ul>
<li>The EMR records visits and OASIS assessments.</li>
<li>A remote monitoring platform logs vitals between visits.</li>
<li>A predictive analytics tool recalculates hospitalization risk scores.</li>
<li>A wound care system captures healing progression with images.</li>
<li>An ambient documentation tool transcribes clinical conversations.</li>
<li>An after-hours triage platform logs patient calls.</li>
</ul>
<p>Every platform does its individual job well. Not one of them shows you the others.</p>
<p>The supervising care team managing 20-40 patients has no realistic way to correlate a vital spike on the remote monitoring platform with a risk score jump on the analytics tool and a missed visit in the EMR, because those three events exist in three separate systems, behind three separate logins, reviewed by three different people on three different timelines!</p>
<p>Check out how individual systems perform their individual roles in the care workflow:</p>
<p><img loading="lazy" decoding="async" class="alignleft wp-image-15350 size-full" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2026/05/How-individual-systems-perform-their-individual-roles-in-the-care-workflow.png" alt="how individual systems perform their individual roles in the care workflow" width="870" height="546" srcset="https://inferenz.ai/wp-content/uploads/2026/05/How-individual-systems-perform-their-individual-roles-in-the-care-workflow.png 870w, https://inferenz.ai/wp-content/uploads/2026/05/How-individual-systems-perform-their-individual-roles-in-the-care-workflow-300x188.png 300w, https://inferenz.ai/wp-content/uploads/2026/05/How-individual-systems-perform-their-individual-roles-in-the-care-workflow-768x482.png 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></p>
<p>The clinical pattern that would predict the next hospitalization is fully present in the data. It just cannot be read simultaneously.</p>
<h2>What Clinical Fragmentation Actually Costs Home Health Agencies</h2>
<p>This is where the stakes become concrete.</p>
<h3>On patient outcomes</h3>
<p>Hospital readmissions remain a major Medicare quality and cost concern, with CMS continuing to tie reimbursement penalties directly to <a href="https://www.cms.gov/medicare/quality/value-based-programs/hospital-readmissions">excess 30-day readmission performance</a>. In home health specifically, the deterioration signals that precede those hospitalizations: weight gain trends, rising vital thresholds, declining ADL scores, missed visits, are almost always present in clinical systems days before the ER visit.</p>
<h3>On Medicare revenue</h3>
<p>A 5% HHVBP payment swing equals $250,000 in annual revenue impact for a $5 million agency. That score is determined by 2024 performance data being calculated right now, as per expanded model. For most agencies, that performance data has never existed in a single unified view. The quality measures driving the score, including Preventable Hospitalization, Discharge Function Score, Discharge to Community, and Medication Management, are each shaped by whether care teams can see patient trajectory across systems in real time.</p>
<h3>On workforce retention</h3>
<p><a href="https://www.hhaexchange.com/blog/recruiting-and-retaining-caregivers">Caregiver turnover sits at 75% annually,</a>a staggering number! Nurses report spending up to two hours per shift navigating disconnected systems to assemble clinical context that should take two minutes. Documentation burden is a structural driver of attrition, not a cultural one. Reducing the time a clinician spends chasing information across platforms is a retention investment, not a workflow convenience.</p>
<h2>What a Unified Patient Timeline Looks Like in Practice</h2>
<p>Before describing how the Patient 360 Journey works technically, it helps to see what changes on a clinical level.</p>
<p>A supervising RN opens a single patient record. Without logging into anything else, she sees:</p>
<ul>
<li><strong>Tuesday:</strong> Blood pressure 158/94, threshold exceeded, flagged moderate severity</li>
<li><strong>Tuesday:</strong> Patient survey reports increased fatigue and mild ankle swelling</li>
<li><strong>Three days prior:</strong> Hospitalization risk score elevated from 38 to 59, contributing factors flagged</li>
<li><strong>Four days prior:</strong> Diuretic dose increased per physician order</li>
<li><strong>Five days prior:</strong> RN visit completed, weight 3.2 lbs above baseline, physician notified</li>
<li><strong>Seven days prior:</strong> Start of Care, primary diagnosis CHF exacerbation</li>
</ul>
<p>That sequence tells a complete clinical story. Rising weight. Medication adjustment. Risk score climbing. Fatigue worsening. Blood pressure spiking. The pattern is unmistakable when all events appear in order on one screen. Without a unified timeline, those same events sit across three platforms, reviewed by different people, connected by nobody.</p>
<p>This is what the Patient 360 Journey makes possible, and it is built entirely from data the organization was already generating. And then the Next Best Action Agent takes it further. It uses the visibility with a recommended next step attached. The right action, for the right patient, delivered to the right person before the pattern becomes a crisis. And it is built entirely from data the organization was already generating.</p>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignleft wp-image-15351 size-full" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2026/05/CTA.jpg" alt="Book a demo CTA" width="870" height="235" srcset="https://inferenz.ai/wp-content/uploads/2026/05/CTA.jpg 870w, https://inferenz.ai/wp-content/uploads/2026/05/CTA-300x81.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/05/CTA-768x207.jpg 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></a></p>
<h2>The Patient 360 Journey and the Next Best Action Agent: How They Work in Four Steps</h2>
<p><img loading="lazy" decoding="async" class="alignleft wp-image-15352 size-full" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2026/05/The-360-Patient-Journey-and-the-Next-Best-Action-Agent-How-They-Work-in-Four-Steps.png" alt="The 360 Patient Journey and the Next Best Action Agent: How They Work in Four Steps" width="870" height="396" srcset="https://inferenz.ai/wp-content/uploads/2026/05/The-360-Patient-Journey-and-the-Next-Best-Action-Agent-How-They-Work-in-Four-Steps.png 870w, https://inferenz.ai/wp-content/uploads/2026/05/The-360-Patient-Journey-and-the-Next-Best-Action-Agent-How-They-Work-in-Four-Steps-300x137.png 300w, https://inferenz.ai/wp-content/uploads/2026/05/The-360-Patient-Journey-and-the-Next-Best-Action-Agent-How-They-Work-in-Four-Steps-768x350.png 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></p>
<h3>Step 1: Centralized Data Warehouse and Master Patient Index</h3>
<p><strong>What it solves:</strong> The same patient carries a different identifier in every system. A medical record number in the EMR. A device ID in remote monitoring. A Medicare beneficiary number in the analytics platform.</p>
<p><strong>How it works:</strong> The Master Patient Index resolves every identifier, including name, date of birth, Medicare ID, and address, into one canonical patient record using probabilistic matching. One patient. One record. Across every system the organization runs.</p>
<p><strong>Why it matters:</strong> Without identity resolution at this level, any downstream unification of clinical data is built on an unreliable foundation. Events get misassigned. Timelines become partial. Clinical decisions get made on incomplete records. The MPI is what makes everything that follows trustworthy.</p>
<h3>Step 2: Standardized Patient Event Model</h3>
<p><strong>What it solves:</strong> Every clinical platform stores data in its own schema, its own timestamp format, its own taxonomy. A vital alert from a remote monitoring platform looks nothing like an OASIS completion from an EMR or a risk score update from a predictive analytics tool.</p>
<p><strong>How it works:</strong> Every clinical event from every connected system gets converted into a single standardized structure: event type, timestamp, source system, clinical status, payload summary, and linked events. The care team does not log into six systems to understand one patient. The data arrives already translated into a common language.</p>
<p><strong>Why it matters:</strong>For example, six systems with six formats produce six incomplete pictures. One standardized event model produces a complete one.</p>
<h3>Step 3: Unified Event Timeline</h3>
<p><strong>What it solves:</strong> Even with data normalized, clinical teams need a way to see the full patient story in sequence, not as a database export.</p>
<p><strong>How it works:</strong> Every normalized event displays in reverse chronological order on a single interface, flagged by severity, color-coded by source system, with linked event relationships visible briefly. The care team sees the complete longitudinal patient journey, from vital spikes and risk score changes to missed visits, wound progression, and after-hours calls, together and in the order they happened.</p>
<p><strong>Why it matters:</strong> Patterns are only visible in sequence. The CHF patient whose weight gain, diuretic adjustment, risk score elevation, and vital spike appear as individual data points across three systems looks like four separate mild concerns. On a single unified timeline, they look like what they are: a hospitalization building over five days.</p>
<h3>Step 4: AI Recommendation Engine and Next Best Action Agent</h3>
<p><strong>What it solves:</strong> A unified timeline shows what happened. The Next Best Action Agent tells care teams what to do about it.</p>
<p><strong>How it works:</strong> The AI Recommendation Engine reads the complete patient timeline and delivers a specific, prioritized recommended action to the right care team member at the right moment. It surfaces patient summaries, risk drivers, and recommended action plans across every risk level, not just critical cases. The right nurse gets the right instruction automatically: schedule a visit today, escalate to the supervisory RN, request reauthorization before the unit gap widens.</p>
<p><strong>Why it matters:</strong> Most clinical AI tools produce dashboards that require interpretation. The Next Best Action Agent produces decisions. There is a meaningful operational difference between a platform that shows a rising risk score and one that tells a specific person to make a specific call within the next four hours.</p>
<h2>How Caregence Connects the Intelligence Layer to Clinical Workflows</h2>
<p>The Next Best Action Agent runs on Caregence, <a href="https://inferenz.ai/healthcare-solutions/caregence-platform/">Inferenz&#8217;s agentic AI platform</a> built specifically for home health and hospice organizations. Caregence connects to existing EMR, payer, scheduling, EVV, and RCM systems without requiring agencies to replace a single platform they already use.</p>
<p>It provides the workflow infrastructure for deploying custom AI agents on top of unified patient data, including the Next Best Action Agent, with built-in governance, role-based access, and audit-ready communication tracking.</p>
<p>Think of Caregence as the operating system for proactive care. The Patient 360 Journey provides the unified data foundation for visibility. It is based on Caregence that provides the AI agents that act on it, including the Next Best Action Agent.</p>
<h2>The Measurable Impact: From Data Visibility to HHVBP Performance</h2>
<p>Inferenz&#8217;s internal assessment of the Patient 360 Journey and Next Best Action Agent against the full HHVBP measure set found that this four-step process addresses up to 63% of HHVBP quality metrics directly.</p>
<p>The measures most influenced:</p>
<p><img loading="lazy" decoding="async" class="alignleft wp-image-15353 size-full" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2026/05/The-Measurable-Impact-From-Data-Visibility-to-HHVBP-Performance.png" alt="The Measurable Impact: From Data Visibility to HHVBP Performance" width="870" height="546" srcset="https://inferenz.ai/wp-content/uploads/2026/05/The-Measurable-Impact-From-Data-Visibility-to-HHVBP-Performance.png 870w, https://inferenz.ai/wp-content/uploads/2026/05/The-Measurable-Impact-From-Data-Visibility-to-HHVBP-Performance-300x188.png 300w, https://inferenz.ai/wp-content/uploads/2026/05/The-Measurable-Impact-From-Data-Visibility-to-HHVBP-Performance-768x482.png 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></p>
<p>The agencies that improve HHVBP scores in 2026 will not do it by changing clinical protocols. They will do it by making existing clinical data visible in sequence, in context, and at the moment when action can still change the outcome.</p>
<h2>The Bottom Line</h2>
<p>Home health and hospice organizations are not data-poor. They are data-fragmented. Every signal needed to prevent the next hospitalization, protect HHVBP reimbursement, reduce documentation burden, and demonstrate outcomes to payers is already being generated inside the organization.</p>
<p>The Patient 360 Journey makes that data readable. Caregence makes it actionable. The Next Best Action Agent makes sure the right person acts on it before the window for intervention closes.</p>
<p>This is what Data to AI to ROI looks like in home health and hospice, built by Inferenz for organizations that cannot afford to keep losing $250,000 on a visibility problem they already have the data to solve.</p>
<h2>Frequently Asked Questions</h2>
<p>The post <a href="https://inferenz.ai/blogs/the-home-health-data-visibility-problem-and-the-ai-agents-that-you-need/">The Home Health Data Visibility Problem and the AI Agents that you Need</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Maximizing Speed, Revenue &#038; Insights with the Right Data Warehouse Design </title>
		<link>https://inferenz.ai/blogs/maximizing-speed-revenue-insights-with-the-right-data-warehouse-design/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Thu, 26 Feb 2026 13:05:00 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://proj.leo9studio.in/projects/inferenz-wp/?p=13967</guid>

					<description><![CDATA[<p>Data warehouse design decides how fast your teams get answers, how much they trust the numbers, and how easily you can scale analytics and AI. This guide breaks down architecture approaches, schema options, and implementation patterns, with clear “use when” guidance for each. </p>
<p>The post <a href="https://inferenz.ai/blogs/maximizing-speed-revenue-insights-with-the-right-data-warehouse-design/">Maximizing Speed, Revenue &amp; Insights with the Right Data Warehouse Design </a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<h2 id="summary" class="wp-block-heading"><strong>Summary</strong></h2>



<p class="wp-block-paragraph">Data warehouse design decides how fast your teams get answers, how much they trust the numbers, and how easily you can scale analytics and AI. This guide breaks down architecture approaches, schema options, and implementation patterns, with clear “use when” guidance for each. </p>



<h2 id="introduction-understanding-data-warehouse-designs" class="wp-block-heading">Introduction: Understanding Data Warehouse Designs </h2>



<p class="wp-block-paragraph">In today’s data-driven world, organizations rely on <strong>data warehouses</strong> to consolidate, organize, and analyze massive volumes of information. But building a data warehouse is not just about storing data – it’s about <strong>designing it in a way that maximizes speed, accuracy, and business value</strong>. </p>



<p class="wp-block-paragraph">A <strong>data warehouse design</strong> determines how data is structured, stored, and accessed. It affects everything from <strong>query performance</strong> to <strong>reporting accuracy</strong>, <strong>machine learning capabilities</strong>, and <strong>regulatory compliance</strong>.  </p>



<p class="wp-block-paragraph">Choosing the right design is crucial because a poorly designed warehouse can slow analytics, increase costs, and lead to incorrect business decisions. </p>



<h2 id="why-data-warehouse-design-matters" class="wp-block-heading">Why Data Warehouse Design Matters </h2>



<ul class="wp-block-list">
<li><strong>Performance:</strong> Ensures queries run quickly, enabling real-time dashboards and faster decision-making. </li>



<li><strong>Scalability:</strong> Supports data growth without costly re-engineering. </li>



<li><strong>Data Quality &amp; Governance:</strong> Reduces redundancy, ensures consistency, and provides audit traceability. </li>



<li><strong>Business Alignment:</strong> Reflects how the business measures success, making analytics intuitive for end-users. </li>
</ul>



<p class="wp-block-paragraph">The following designs apply to organizations that provide data in batches. Details on warehouse design for organizations that provide real-time data, will be covered separately. </p>



<h2 id="simple-data-warehouse-architecture-diagram-3-layer-view" class="wp-block-heading">Simple Data Warehouse Architecture Diagram (3-Layer View) </h2>



<figure class="wp-block-gallery has-nested-images columns-default is-cropped wp-block-gallery-1 is-layout-flex wp-block-gallery-is-layout-flex">
<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="637" height="1024" data-id="14925" class="wp-image-14925" src="https://inferenz.ai/wp-content/uploads/2026/04/Simple-Data-Warehouse-Architecture-Diagram-3-Layer-View-637x1024.jpg" alt="Simple Data Warehouse Architecture Diagram (3-Layer View) " srcset="https://inferenz.ai/wp-content/uploads/2026/04/Simple-Data-Warehouse-Architecture-Diagram-3-Layer-View-637x1024.jpg 637w, https://inferenz.ai/wp-content/uploads/2026/04/Simple-Data-Warehouse-Architecture-Diagram-3-Layer-View-187x300.jpg 187w, https://inferenz.ai/wp-content/uploads/2026/04/Simple-Data-Warehouse-Architecture-Diagram-3-Layer-View-768x1234.jpg 768w, https://inferenz.ai/wp-content/uploads/2026/04/Simple-Data-Warehouse-Architecture-Diagram-3-Layer-View.jpg 871w" sizes="auto, (max-width: 637px) 100vw, 637px" /></figure>
</figure>



<p class="wp-block-paragraph"><strong>Source systems</strong> <br />ERP, CRM, product apps, files, APIs, event streams </p>



<p class="wp-block-paragraph"><strong>Ingestion and integration</strong> <br />ETL or ELT, CDC, data quality checks, standardization </p>



<p class="wp-block-paragraph"><strong>Warehouse and modeling layers</strong> <br />Architecture approach (Kimball, Inmon, Data Vault, Anchor) <br />Schema design (star, snowflake, galaxy, 3NF) <br />Implementation patterns (wide tables, aggregates, hybrid) </p>



<p class="wp-block-paragraph"><strong>Consumption</strong> <br />BI tools, dashboards, ad-hoc queries, ML workflows</p>



<h2 id="data-warehouse-architecture-design-approaches" class="wp-block-heading">Data Warehouse Architecture / Design Approaches </h2>



<p class="wp-block-paragraph">Data warehouse architecture defines the <strong>overall strategy and methodology</strong> for building a data warehouse, guiding how data is collected, integrated, stored, and accessed for analysis. Unlike individual schema designs that focus on table structures, these approaches provide a <strong>high-level blueprint</strong> for enterprise data management and analytics. </p>



<h3 class="wp-block-heading"><strong>Kimball Dimensional Modeling</strong> </h3>



<p class="wp-block-paragraph">Kimball focuses on building dimensional models around business processes, often as data marts that roll up into a broader analytical layer. It is popular because it is easy to understand and fast for BI. </p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="692" height="1024" class="wp-image-15001" src="https://inferenz.ai/wp-content/uploads/2026/02/Kimball-Dimensional-Modeling-1-692x1024.jpg" alt="Kimball Dimensional Modeling" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Kimball-Dimensional-Modeling-1-692x1024.jpg 692w, https://inferenz.ai/wp-content/uploads/2026/02/Kimball-Dimensional-Modeling-1-203x300.jpg 203w, https://inferenz.ai/wp-content/uploads/2026/02/Kimball-Dimensional-Modeling-1-768x1137.jpg 768w, https://inferenz.ai/wp-content/uploads/2026/02/Kimball-Dimensional-Modeling-1.jpg 870w" sizes="auto, (max-width: 692px) 100vw, 692px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>Business users need intuitive reporting quickly </li>



<li>Requirements are stable and well understood </li>



<li>You want incremental delivery with visible wins </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<ul class="wp-block-list">
<li>BI dashboards, finance and revenue reporting, sales and marketing analytics </li>
</ul>



<p class="wp-block-paragraph"><strong>Typical impact</strong> </p>



<ul class="wp-block-list">
<li>Faster time to value, strong user adoption, simpler reporting model </li>
</ul>



<p class="wp-block-paragraph"><strong>Example scenario</strong> <br />Marketing needs campaign performance dashboards quickly. Kimball supports focused data marts, conformed dimensions, and fast reporting delivery. </p>



<h3 class="wp-block-heading"><strong>Inmon top-down approach (enterprise-first EDW)</strong> </h3>



<p class="wp-block-paragraph">Inmon starts with a centralized enterprise data warehouse, usually in normalized 3NF structures. Data marts are derived later for performance and ease of reporting. It takes longer to build but supports consistent enterprise definitions. </p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="574" height="1024" class="wp-image-14985" src="https://inferenz.ai/wp-content/uploads/2026/02/Inmon-top-down-approach-enterprise-first-EDW-574x1024.jpg" alt="Inmon top-down approach (enterprise-first EDW)" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Inmon-top-down-approach-enterprise-first-EDW-574x1024.jpg 574w, https://inferenz.ai/wp-content/uploads/2026/02/Inmon-top-down-approach-enterprise-first-EDW-168x300.jpg 168w, https://inferenz.ai/wp-content/uploads/2026/02/Inmon-top-down-approach-enterprise-first-EDW-768x1369.jpg 768w, https://inferenz.ai/wp-content/uploads/2026/02/Inmon-top-down-approach-enterprise-first-EDW-861x1536.jpg 861w, https://inferenz.ai/wp-content/uploads/2026/02/Inmon-top-down-approach-enterprise-first-EDW.jpg 871w" sizes="auto, (max-width: 574px) 100vw, 574px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>A single version of truth is required across functions and regions </li>



<li>Governance and standardization are priorities </li>



<li>Integration across many systems is complex </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<ul class="wp-block-list">
<li>Large enterprises with strict KPI consistency and governance needs </li>
</ul>



<p class="wp-block-paragraph"><strong>Typical impact</strong> </p>



<ul class="wp-block-list">
<li>Higher trust in metrics, stronger control, better enterprise alignment </li>
</ul>



<p class="wp-block-paragraph"><strong>Example scenario</strong> <br />A global company needs standardized KPIs across regions. Inmon supports centralized definitions and reduces conflicting reports. </p>



<h3 class="wp-block-heading"><strong>Data Vault modeling (scalable and auditable)</strong> </h3>



<p class="wp-block-paragraph">Data Vault organizes data into <strong>Hubs</strong> (business keys), <strong>Links</strong> (relationships), and <strong>Satellites</strong> (descriptive history). It separates raw ingestion from business logic, which helps with change, traceability, and long-term integration. </p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="663" height="1024" class="wp-image-14987" src="https://inferenz.ai/wp-content/uploads/2026/02/Data-Vault-modeling-scalable-and-auditable-663x1024.jpg" alt="Data Vault modeling (scalable and auditable)" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Data-Vault-modeling-scalable-and-auditable-663x1024.jpg 663w, https://inferenz.ai/wp-content/uploads/2026/02/Data-Vault-modeling-scalable-and-auditable-194x300.jpg 194w, https://inferenz.ai/wp-content/uploads/2026/02/Data-Vault-modeling-scalable-and-auditable-768x1186.jpg 768w, https://inferenz.ai/wp-content/uploads/2026/02/Data-Vault-modeling-scalable-and-auditable.jpg 871w" sizes="auto, (max-width: 663px) 100vw, 663px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>Source systems change often </li>



<li>Historical tracking and auditability matter </li>



<li>You expect new domains and sources over time </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<ul class="wp-block-list">
<li>Telecom, finance, insurance, regulated industries, complex enterprise integration </li>
</ul>



<p class="wp-block-paragraph"><strong>Typical impact</strong> </p>



<ul class="wp-block-list">
<li>Faster onboarding of sources, fewer breakages from schema drift, stronger lineage </li>
</ul>



<p class="wp-block-paragraph"><strong>Example scenario</strong> <br />A telecom adds new products and pricing models often. Data Vault reduces the blast radius of change and keeps history intact. </p>



<h3 class="wp-block-heading"><strong>Anchor modeling (high adaptability in a normalized style)</strong> </h3>



<p class="wp-block-paragraph">Anchor modeling uses <strong>Anchors</strong> (core entities), <strong>Attributes</strong>, and <strong>Ties</strong> (relationships). It is designed for frequent change. You can add new attributes without redesigning large parts of the model. </p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="618" height="1024" class="wp-image-14988" src="https://inferenz.ai/wp-content/uploads/2026/02/Anchor-modeling-high-adaptability-in-a-normalized-style-618x1024.jpg" alt="Anchor modeling (high adaptability in a normalized style)" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Anchor-modeling-high-adaptability-in-a-normalized-style-618x1024.jpg 618w, https://inferenz.ai/wp-content/uploads/2026/02/Anchor-modeling-high-adaptability-in-a-normalized-style-181x300.jpg 181w, https://inferenz.ai/wp-content/uploads/2026/02/Anchor-modeling-high-adaptability-in-a-normalized-style-768x1272.jpg 768w, https://inferenz.ai/wp-content/uploads/2026/02/Anchor-modeling-high-adaptability-in-a-normalized-style.jpg 871w" sizes="auto, (max-width: 618px) 100vw, 618px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>Business attributes and rules change frequently </li>



<li>You need flexibility without major table redesign </li>



<li>You want a long-lived model that evolves with the business </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<ul class="wp-block-list">
<li>Fast-changing SaaS environments and evolving product analytics needs </li>
</ul>



<p class="wp-block-paragraph"><strong>Typical impact</strong> </p>



<ul class="wp-block-list">
<li>Less rework, easier schema evolution, better maintainability </li>
</ul>



<p class="wp-block-paragraph"><strong>Example scenario</strong> <br />A SaaS business keeps adding customer attributes. Anchor modeling supports this without downtime-heavy redesigns. </p>



<figure class="wp-block-image size-full"><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" width="871" height="235" class="wp-image-14990" src="https://inferenz.ai/wp-content/uploads/2026/02/CTA-1-1.jpg" alt="CTA" srcset="https://inferenz.ai/wp-content/uploads/2026/02/CTA-1-1.jpg 871w, https://inferenz.ai/wp-content/uploads/2026/02/CTA-1-1-300x81.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/02/CTA-1-1-768x207.jpg 768w" sizes="auto, (max-width: 871px) 100vw, 871px" /></a></figure>



<h2 id="schema-designs-logical-and-physical-models" class="wp-block-heading">Schema designs: logical and physical models </h2>



<p class="wp-block-paragraph">Schemas define how tables are structured. They affect join patterns, usability, and performance. </p>



<h3 class="wp-block-heading"><strong>Star schema </strong></h3>



<p class="wp-block-paragraph">The star schema is a central fact table that connects to denormalized dimensions. It is widely used because it is fast and easy to query. </p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="871" height="377" class="wp-image-14991" src="https://inferenz.ai/wp-content/uploads/2026/02/Star-schema.jpg" alt="Star schema" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Star-schema.jpg 871w, https://inferenz.ai/wp-content/uploads/2026/02/Star-schema-300x130.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/02/Star-schema-768x332.jpg 768w" sizes="auto, (max-width: 871px) 100vw, 871px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>You want fast BI and simple reporting </li>



<li>Many users run ad-hoc analysis </li>



<li>Business teams need clear dimensions and metrics </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<p class="wp-block-paragraph">Dashboards, KPI reporting, analytics that depend on speed </p>



<h3 class="wp-block-heading"><strong>Snowflake schema </strong></h3>



<p class="wp-block-paragraph">Dimensions are normalized into sub-tables, often to manage hierarchies and reduce redundancy. It can save storage but adds joins. </p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="871" height="451" class="wp-image-14992" src="https://inferenz.ai/wp-content/uploads/2026/02/Snowflake-schema.jpg" alt="Snowflake schema" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Snowflake-schema.jpg 871w, https://inferenz.ai/wp-content/uploads/2026/02/Snowflake-schema-300x155.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/02/Snowflake-schema-768x398.jpg 768w" sizes="auto, (max-width: 871px) 100vw, 871px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>Dimension hierarchies are complex </li>



<li>Storage efficiency matters </li>



<li>Slightly slower queries are acceptable </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<p class="wp-block-paragraph">Large product catalogs, structured hierarchies, domains with frequent hierarchy updates </p>



<h3 class="wp-block-heading"><strong>Galaxy schema (fact constellation)</strong> </h3>



<p class="wp-block-paragraph">Multiple fact tables share dimension tables. It supports cross-process analytics across domains like orders, shipments, returns, and inventory. </p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="871" height="378" class="wp-image-14993" src="https://inferenz.ai/wp-content/uploads/2026/02/Galaxy-schema.jpg" alt="Galaxy schema (fact constellation)" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Galaxy-schema.jpg 871w, https://inferenz.ai/wp-content/uploads/2026/02/Galaxy-schema-300x130.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/02/Galaxy-schema-768x333.jpg 768w" sizes="auto, (max-width: 871px) 100vw, 871px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>You need analysis across multiple business processes </li>



<li>Shared dimensions create enterprise views of the customer or product </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<p class="wp-block-paragraph">E-commerce, supply chain, end-to-end customer journey analytics </p>



<h3 class="wp-block-heading"><strong>Normalized 3NF enterprise warehouse </strong></h3>



<p class="wp-block-paragraph">Highly normalized tables reduce redundancy and enforce integrity. It is strong for integration and governance, but reporting queries can be slower without downstream marts. </p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="871" height="302" class="wp-image-14994" src="https://inferenz.ai/wp-content/uploads/2026/02/Normalized-3NF-enterprise-warehouse.jpg" alt="Normalized 3NF enterprise warehouse" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Normalized-3NF-enterprise-warehouse.jpg 871w, https://inferenz.ai/wp-content/uploads/2026/02/Normalized-3NF-enterprise-warehouse-300x104.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/02/Normalized-3NF-enterprise-warehouse-768x266.jpg 768w" sizes="auto, (max-width: 871px) 100vw, 871px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>The warehouse is a system of record </li>



<li>Audit and regulatory demands are high </li>



<li>Integration consistency matters more than reporting speed </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<p class="wp-block-paragraph">Enterprise integration layer, regulated domains, “one source of truth” requirements </p>



<h2 id="physical-implementation-patterns" class="wp-block-heading">Physical implementation patterns </h2>



<p class="wp-block-paragraph">These patterns influence performance and cost once architecture and schemas are chosen. </p>



<h3 class="wp-block-heading"><strong>Wide tables </strong></h3>



<p class="wp-block-paragraph">Wide tables store facts and useful attributes together in a denormalized structure. They reduce joins and speed up analytics and ML feature use.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="870" height="301" class="wp-image-14995" src="https://inferenz.ai/wp-content/uploads/2026/02/Wide-tables.png" alt="Wide tables " srcset="https://inferenz.ai/wp-content/uploads/2026/02/Wide-tables.png 870w, https://inferenz.ai/wp-content/uploads/2026/02/Wide-tables-300x104.png 300w, https://inferenz.ai/wp-content/uploads/2026/02/Wide-tables-768x266.png 768w" sizes="auto, (max-width: 870px) 100vw, 870px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>ML feature pipelines suffer from join complexity </li>



<li>Query speed is more important than storage </li>



<li>Data models are stable enough for denormalization </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<ul class="wp-block-list">
<li>AI feature stores, customer 360 analytics, experimentation analytics </li>
</ul>



<h3 class="wp-block-heading"><strong>Hybrid designs </strong></h3>



<p class="wp-block-paragraph">Hybrid designs mix approaches and optimize each layer for its job. A common pattern is: raw integration layer (often Data Vault), then dimensional marts for BI, then wide tables for ML and performance-heavy use cases.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="871" height="564" class="wp-image-14996" src="https://inferenz.ai/wp-content/uploads/2026/02/Hybrid-designs.jpg" alt="Hybrid designs" srcset="https://inferenz.ai/wp-content/uploads/2026/02/Hybrid-designs.jpg 871w, https://inferenz.ai/wp-content/uploads/2026/02/Hybrid-designs-300x194.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/02/Hybrid-designs-768x497.jpg 768w" sizes="auto, (max-width: 871px) 100vw, 871px" /></figure>



<p class="wp-block-paragraph"><strong>Use when</strong> </p>



<ul class="wp-block-list">
<li>You support BI, advanced analytics, and ML together </li>



<li>Workloads differ by team and tool </li>



<li>You want both governance and speed </li>
</ul>



<p class="wp-block-paragraph"><strong>Best fit</strong> </p>



<ul class="wp-block-list">
<li>Modern enterprise data platforms where one model cannot satisfy every use case </li>
</ul>



<h2 id="practical-selection-guide" class="wp-block-heading">Practical selection guide </h2>



<ul class="wp-block-list">
<li><strong>Fast reporting and quick wins:</strong> Kimball + star schema </li>



<li><strong>Enterprise consistency and governance:</strong> Inmon or 3NF EDW feeding marts </li>



<li><strong>Frequent source change and deep audit needs:</strong> Data Vault </li>



<li><strong>Rapidly evolving attributes and long-term flexibility:</strong> Anchor modeling </li>



<li><strong>Cross-domain process analytics:</strong> Galaxy schema </li>



<li><strong>Performance-heavy analytics and ML features:</strong> Wide tables </li>



<li><strong>Mixed workloads across BI and AI:</strong> Hybrid layered approach </li>
</ul>



<h2 id="conclusion" class="wp-block-heading">Conclusion </h2>



<p class="wp-block-paragraph">The best <strong>data warehouse design</strong> is the one that fits your business reality, not the one that looks best on a whiteboard. Every architecture choice shapes what happens downstream: dashboard speed, reporting trust, integration effort, governance strength, and how ready your teams are for advanced analytics and AI. </p>



<p class="wp-block-paragraph">For most U.S. enterprises, the smartest path is to separate concerns. Use a strong <strong>data warehouse architecture</strong> for integration and traceability, choose the right <strong>data warehouse schema design</strong> for reporting, and apply performance patterns like <strong>wide table design</strong> only where they make sense. In many environments, that naturally leads to a <strong>hybrid data warehouse architecture</strong>, where <strong>Data Vault modeling</strong> supports scalable ingestion, <strong>Kimball dimensional modeling</strong> powers BI adoption, and curated layers enable ML without breaking reporting. </p>



<p class="wp-block-paragraph">Whether you choose the <strong>Inmon approach</strong>, a pure dimensional strategy, or a layered model, the goal stays the same: reduce friction between data teams and decision-makers. When the design is right, analytics becomes faster, costs become predictable, and the warehouse becomes a stable foundation for growth, modernization, and AI-driven outcomes. </p>



<figure class="wp-block-image size-full"><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" width="871" height="236" class="wp-image-14997" src="https://inferenz.ai/wp-content/uploads/2026/02/CTA-2.jpg" alt="CTA 2" srcset="https://inferenz.ai/wp-content/uploads/2026/02/CTA-2.jpg 871w, https://inferenz.ai/wp-content/uploads/2026/02/CTA-2-300x81.jpg 300w, https://inferenz.ai/wp-content/uploads/2026/02/CTA-2-768x208.jpg 768w" sizes="auto, (max-width: 871px) 100vw, 871px" /></a></figure>



<h2 id="frequently-asked-questions" class="wp-block-heading">Frequently Asked Questions</h2>



<ol class="wp-block-list" start="1"></ol>
<p>The post <a href="https://inferenz.ai/blogs/maximizing-speed-revenue-insights-with-the-right-data-warehouse-design/">Maximizing Speed, Revenue &amp; Insights with the Right Data Warehouse Design </a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>FinOps in Real-World Practice: Transforming Cloud Spend into Strategic Value</title>
		<link>https://inferenz.ai/blogs/finops-in-real-world-practice-transforming-cloud-spend-into-strategic-value/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Wed, 07 Jan 2026 11:08:17 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://inferenz.ai/?p=12412</guid>

					<description><![CDATA[<p>As cloud adoption grows in fintech, cloud cost management becomes harder because usage and pricing shift every hour. FinOps helps teams link spend to real outcomes like cost per transaction, fraud checks, and feature delivery.</p>
<p>The post <a href="https://inferenz.ai/blogs/finops-in-real-world-practice-transforming-cloud-spend-into-strategic-value/">FinOps in Real-World Practice: Transforming Cloud Spend into Strategic Value</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<h2><span class="TextRun SCXW37838156 BCX0" lang="EN-IN" xml:lang="EN-IN" data-contrast="auto"><span class="NormalTextRun SCXW37838156 BCX0">Summary</span></span></h2>
<p><i><span style="font-weight: 400;">As cloud adoption grows in fintech, cloud cost management becomes harder because usage and pricing shift every hour. FinOps helps teams link spend to real outcomes like cost per transaction, fraud checks, and feature delivery. Learn how fintech teams apply FinOps in daily operations, using tagging, visibility, forecasting, and automation to turn cloud spend into strategic value.</span></i><img loading="lazy" decoding="async" class="alignleft wp-image-12416 size-full" style="width: 100%; display: block; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2026/01/Cloud-Spend-to-Strategic-Value-with-FinOps.jpg" alt="Cloud spend to strategic value with FinOps" width="1440" height="1029" /></p>
<h2>Introduction</h2>
<p><span style="font-weight: 400;">Cloud makes fintech faster. Teams can ship features quickly, scale during peak transaction windows, and run analytics without buying hardware. </span></p>
<p><span style="font-weight: 400;">The catch is simple: consumption pricing turns every new workload into a variable cost line. And in fintech, workloads spike for reasons that feel “business as usual” such as payout cycles, fraud bursts, seasonal lending, or a partner API change.</span></p>
<p><span style="font-weight: 400;">FinOps exists to keep that variability from becoming chaos. The FinOps Foundation defines FinOps as an </span><a href="https://www.finops.org/introduction/what-is-finops/"><span style="font-weight: 400;">operational framework and cultural practice</span></a><span style="font-weight: 400;"> that maximizes business value from cloud and technology through timely, data-driven decisions and shared financial accountability across engineering, finance, and business teams. </span></p>
<p><span style="font-weight: 400;">This guide shows what FinOps looks like when you apply it day to day in fintech environments, where speed, governance, and predictability matter at the same time.</span></p>
<h2><span style="font-weight: 400;">Why fintech teams feel cloud cost pressure sooner</span></h2>
<p><span style="font-weight: 400;">Fintech cloud usage tends to concentrate in a few expensive areas:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Always-on customer experiences</b><span style="font-weight: 400;">: low-latency apps, APIs, identity, and observability.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Risk and fraud analytics</b><span style="font-weight: 400;">: streaming, feature stores, model training, and bursty compute.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Data platforms</b><span style="font-weight: 400;">: warehouses and lakehouses that grow quietly with retention, audit, and regulatory needs.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Security controls</b><span style="font-weight: 400;">: logging, monitoring, scanning, and encryption overhead that is necessary, but rarely “free.”</span></li>
</ul>
<p><span style="font-weight: 400;">And cloud spend keeps climbing across industries. Gartner forecasts public </span><a href="https://www.gartner.com/en/newsroom/press-releases/2024-11-19-gartner-forecasts-worldwide-public-cloud-end-user-spending-to-total-723-billion-dollars-in-2025"><span style="font-weight: 400;">cloud end-user spending</span></a><span style="font-weight: 400;"> at </span><b>$723.4B in 2025</b><span style="font-weight: 400;">. </span></p>
<p><span style="font-weight: 400;">So, the question for fintech leaders is rarely “should we spend less?” It’s “how do we spend with intent, and prove it with numbers?”</span></p>
<p><span style="font-weight: 400;">That’s where FinOps becomes a business discipline, not a billing exercise.</span></p>
<p><img loading="lazy" decoding="async" class="alignleft wp-image-12421 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2026/01/Three-phases-of-FinOps-%E2%80%93-Inform-Optimize-and-Operate.jpg" alt="Three phases of FinOps" width="2000" height="1467" /></p>
<h2 style="margin-top: 20px;">FinOps in daily operations: the practices that change outcomes</h2>
<h3><span style="font-weight: 400;">1) Unify teams around shared financial accountability</span></h3>
<p><span style="font-weight: 400;">FinOps works when engineering and finance stop treating cloud cost as someone else’s job. The practical shift looks like this:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Finance gets </span><b>clear ownership views</b><span style="font-weight: 400;">: by product, environment, and business line.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Engineering gets </span><b>fast feedback loops</b><span style="font-weight: 400;">: cost impact is visible before and after a release.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Product and leadership get </span><b>unit economics</b><span style="font-weight: 400;">: cost per transaction, cost per active customer, cost per underwriting decision, cost per fraud check.</span></li>
</ul>
<p><b><i>Example</i></b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Before launching a new real-time payments feature, the platform team reviews expected throughput, storage growth, and observability overhead with finance. They agree on a target unit cost (say, cost per 1,000 transactions) and track it weekly. If unit cost rises, teams investigate whether it came from higher log volume, unbounded retries, or an over-sized compute tier.</span></p>
<p><span style="font-weight: 400;">What Inferenz typically adds here is the operating model: who owns which cost domains, what gets reviewed weekly versus monthly, and how teams turn cost data into decisions without slowing delivery.</span></p>
<h3><span style="font-weight: 400;">2) Make cost visibility usable with tagging, allocation, and clean data</span></h3>
<p><span style="font-weight: 400;">Visibility is more than a dashboard. It’s consistent, trusted allocation that supports action.</span></p>
<p><span style="font-weight: 400;">For fintech teams, a tagging and allocation baseline usually includes:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Product / business line</b></li>
<li style="font-weight: 400;" aria-level="1"><b>Environment</b><span style="font-weight: 400;"> (prod, staging, dev)</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Cost center</b></li>
<li style="font-weight: 400;" aria-level="1"><b>Workload type</b><span style="font-weight: 400;"> (API, batch, streaming, ML training, BI)</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Data classification</b><span style="font-weight: 400;"> (helps align cost with governance and audit needs)</span></li>
</ul>
<p><span style="font-weight: 400;">Tools such as AWS Cost Explorer and Azure Cost Management help, but they depend on </span><a href="https://learn.microsoft.com/en-us/cloud-computing/finops/overview"><span style="font-weight: 400;">clean tagging and consistent account structure</span></a><span style="font-weight: 400;">.</span></p>
<p><b><i>Quick win that matters:</i></b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Create a “no tag, no launch” gate for production infrastructure as a guardrail that prevents unknown spend from becoming permanent.</span></p>
<p><a href="https://inferenz.ai/blogs/data-quality-and-governance-for-scalable-and-sustainable-growth/"><img loading="lazy" decoding="async" class="alignleft wp-image-12417 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2026/01/CTA-1.gif" alt="Data quality and governance blog" width="1400" height="378" /></a></p>
<h3 style="margin-top: 20px;">3) Shift from month-end surprises to real-time decisions</h3>
<p><span style="font-weight: 400;">FinOps teams operate on short cycles because cloud changes daily. When cost signals arrive a month later, the money is already gone.</span></p>
<p><span style="font-weight: 400;">In real practice, fintech teams do things like:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Auto-shutdown</b><span style="font-weight: 400;"> non-critical environments after hours</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Rightsize compute</b><span style="font-weight: 400;"> based on actual utilization</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Use commitment planning</b><span style="font-weight: 400;"> (Savings Plans, Reserved Instances) where usage is steady</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Move storage</b><span style="font-weight: 400;"> to lower-cost tiers with policy-based lifecycle rules</span></li>
</ul>
<p><span style="font-weight: 400;">FinOps Foundation guidance frames this as a </span><a href="https://www.finops.org/framework/?"><span style="font-weight: 400;">loop across visibility</span></a><span style="font-weight: 400;">, optimization, and operations. </span></p>
<p><b><i>Example</i></b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">A fraud model retrains nightly. The pipeline grew over time and now runs on larger nodes than needed. FinOps flags the change in cost per training run, the data team confirms stable runtime targets, and the platform team applies right-sizing and schedule controls. The end result is predictable spend without weakening detection.</span></p>
<h3><span style="font-weight: 400;">4) Treat forecasting like a product KPI, not a finance exercise</span></h3>
<p><span style="font-weight: 400;">Forecasting is where fintech teams often struggle because demand is real-time and spiky. Still, you can forecast well if you forecast the right thing.</span></p>
<p><span style="font-weight: 400;">Instead of asking, “</span><i><span style="font-weight: 400;">What will AWS bill be next month?”,</span></i><span style="font-weight: 400;"> focus on:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">forecasted </span><b>unit volumes</b><span style="font-weight: 400;"> (transactions, API calls, onboarding checks)</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">expected </span><b>model usage</b><span style="font-weight: 400;"> (training runs, inference calls)</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">the unit cost curve (cost per 1,000 events)</span></li>
</ul>
<p><span style="font-weight: 400;">Then tie cloud spend to those business drivers.</span></p>
<p><a href="https://www.flexera.com/about-us/press-center/new-flexera-report-finds-84-percent-of-organizations-struggle-to-manage-cloud-spend"><span style="font-weight: 400;">Cloud spend management</span></a><span style="font-weight: 400;"> remains a widespread challenge, which makes forecasting discipline a differentiator.</span></p>
<p><span style="font-weight: 400;">Where Inferenz fits: building data pipelines that merge billing exports, usage telemetry, and product metrics so forecasts reflect how the business actually runs, beyond what the invoice says.</span></p>
<h2>How fintech teams scale FinOps by maturity</h2>
<p><img loading="lazy" decoding="async" class="alignleft wp-image-12419 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2026/01/How-fintech-teams-scale-FinOps-by-maturity.jpg" alt="How fintech teams scale FinOps by maturity" width="2000" height="1467" /></p>
<h2>Common roadblocks and how to get past them</h2>
<p><img loading="lazy" decoding="async" class="alignleft wp-image-12422 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2026/01/Three-obstacles-to-scaling-FinOps.jpg" alt="Three obstacles to scaling FinOps" width="1440" height="1029" /></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Resistance from teams</b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Engineers may assume cost controls will slow delivery. Fix that by using automation, clear thresholds, and fast feedback, not manual approvals.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Complex pricing and confusing bills</b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Cloud pricing is hard. The fix is to translate billing into “engineering terms” such as runtime, storage growth, egress, and query patterns.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Inconsistent governance<br />
</b>If tagging rules vary by team, visibility collapses. Standardize the minimum required tags and enforce them with policy.</li>
</ul>
<h2>Recommended practices for sustainable FinOps adoption in fintech</h2>
<p><img loading="lazy" decoding="async" class="alignleft wp-image-12420 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2026/01/Recommended-practices-for-sustainable-FinOps-adoption-in-fintech.jpg" alt="Recommended practices for sustainable FinOps adoption in fintech" width="2000" height="1467" /></p>
<ol>
<li style="font-weight: 400;" aria-level="1"><b>Start with 1 or 2 high-impact domains</b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Common picks: fraud analytics pipeline, core API platform, data warehouse.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Define unit economics everyone understands</b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Cost per transaction, cost per onboarded customer, cost per underwriting decision.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Automate guardrails</b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Idle cleanup, tag enforcement, budget alerts, and anomaly detection.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Make the weekly FinOps review short and decisive</b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Review top cost drivers, anomalies, and planned changes for next week.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Tie spend to business outcomes</b><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">Revenue growth, authorization rates, fraud loss reduction, time-to-ship, or customer experience KPIs.</span></li>
</ol>
<h2><span style="font-weight: 400;">Final thoughts</span></h2>
<p><span style="font-weight: 400;">FinOps becomes valuable in fintech when it connects cloud spend to product reality: usage, risk controls, and customer outcomes. With the right allocation, unit economics, and automation, teams keep speed while making spend predictable and defensible.</span><span style="font-weight: 400;"><br />
</span></p>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignleft wp-image-12418 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2026/01/CTA-2.gif" alt="CTA Contact Us" width="1400" height="378" /></a></p>
<h2><span style="font-weight: 400;">Frequently asked questions</span></h2>
<p>The post <a href="https://inferenz.ai/blogs/finops-in-real-world-practice-transforming-cloud-spend-into-strategic-value/">FinOps in Real-World Practice: Transforming Cloud Spend into Strategic Value</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Data Quality &#038; Governance: The Strategic Blueprint for Sustainable Organizational Success</title>
		<link>https://inferenz.ai/blogs/data-quality-and-governance-for-scalable-and-sustainable-growth/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Wed, 03 Dec 2025 09:30:45 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://inferenz.ai/?p=12281</guid>

					<description><![CDATA[<p>In an era defined by data, organizations are navigating a fundamental paradox: they are data-rich but insight-poor. The sheer volume of information, intended to be a strategic asset for every Fortune 100 contender and nimble startup alike, often becomes a source of complexity and confusion.  Without a structured approach, this asset quickly turns into a [&#8230;]</p>
<p>The post <a href="https://inferenz.ai/blogs/data-quality-and-governance-for-scalable-and-sustainable-growth/">Data Quality &#038; Governance: The Strategic Blueprint for Sustainable Organizational Success</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><span style="font-weight: 400;">In an era defined by data, organizations are navigating a fundamental paradox: they are data-rich but insight-poor. The sheer volume of information, intended to be a strategic asset for every </span><b>Fortune 100</b><span style="font-weight: 400;"> contender and nimble startup alike, often becomes a source of complexity and confusion. </span></p>
<p><span style="font-weight: 400;">Without a structured approach, this asset quickly turns into a liability, leading to flawed strategies, missed opportunities, and eroded trust. The solution is not more data, but better, more reliable data, managed under a coherent strategic framework. This is the essence of data quality and governance: the strategic blueprint for transforming data chaos into a sustainable competitive advantage.</span></p>
<h2><span style="font-weight: 400;">The data imperative: Why trustworthy data is non-negotiable</span></h2>
<p><span style="font-weight: 400;">In today&#8217;s digital economy, every critical business function relies on data. From personalizing a customer journey to optimizing supply chains with </span><b>big data</b><span style="font-weight: 400;"> analytics, the accuracy and reliability of the underlying information dictate the outcome. </span></p>
<p><span style="font-weight: 400;">Poor </span><b>Data Quality</b><span style="font-weight: 400;"> directly translates to poor decision-making, misguided strategies, and inefficient operations. When leadership cannot trust the numbers presented in a </span><b>Business intelligence</b><span style="font-weight: 400;"> dashboard, strategic planning becomes a game of guesswork, and the organization’s ability to respond to market shifts is severely compromised. </span></p>
<p><span style="font-weight: 400;">Trustworthy data is the foundational prerequisite for organizational agility and resilience.</span></p>
<h2><span style="font-weight: 400;">The Promise of AI: unlocking potential through data excellence</span></h2>
<p><span style="font-weight: 400;">AI initiatives promise to change industries. However, AI is not magic; it is a sophisticated consumer of data. </span></p>
<p><b>Machine learning algorithms</b><span style="font-weight: 400;"> are only as effective as the data they are trained on. Biased, incomplete, or inaccurate data leads to flawed models, unreliable predictions, and potentially disastrous business outcomes. A staggering number of AI projects fail to move from pilot to production, not because the algorithms are weak, but because the data foundation is unstable. </span></p>
<p><span style="font-weight: 400;">True </span><b>AI Readiness</b><span style="font-weight: 400;"> begins with a deep commitment to data quality and governance, ensuring that your most advanced initiatives are built on a bedrock of trust.</span></p>
<h2><span style="font-weight: 400;">Setting the stage: Data Quality and Governance as your strategic foundation</span></h2>
<p><span style="font-weight: 400;">Viewing data quality and governance as mere compliance obligations or IT-centric tasks is a critical strategic error. Instead, they must be positioned as the central pillars of an organization&#8217;s data strategy: the keys to why </span><b>Data Quality</b><span style="font-weight: 400;"> and governance drive digital success. A robust governance framework acts as the control system, defining the rules of engagement for all data assets, while a commitment to data quality ensures those assets are fit for purpose. </span></p>
<p><span style="font-weight: 400;">Together, they create an environment where data can be confidently accessed, shared, and leveraged to drive innovation and create tangible business value, forming the strategic blueprint for enduring success.</span></p>
<h2><span style="font-weight: 400;">The indispensable foundation: Unpacking Data Quality and Governance</span></h2>
<p><span style="font-weight: 400;">Before building a data-driven enterprise, leaders must understand the core components of its foundation. </span><b>Data Quality</b><span style="font-weight: 400;"> and data governance are distinct but deeply interconnected disciplines. One cannot succeed without the other. Governance provides the structure, rules, and accountability, while quality represents the tangible, measurable state of the data itself.</span></p>
<h3><span style="font-weight: 400;">Defining Data Quality: dimensions of trust</span><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12288" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2025/12/Defining-Data-Quality-dimensions-of-trust.jpg" alt="" width="1440" height="1029" /></h3>
<p><b>Data Quality</b><span style="font-weight: 400;"> is not a single attribute but a multi-dimensional concept, often defined by standards like </span><b>ISO/IEC 25012</b><span style="font-weight: 400;">. To be considered high-quality, data must meet several key criteria:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Accuracy:</b><span style="font-weight: 400;"> Does the data correctly reflect the real-world object or event it describes?</span></li>
</ul>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Completeness:</b><span style="font-weight: 400;"> Are all the necessary data points present?</span></li>
</ul>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Consistency:</b><span style="font-weight: 400;"> Is the data uniform across different systems and applications?</span></li>
</ul>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Timeliness:</b><span style="font-weight: 400;"> Is the data available when it is needed for analysis and decision-making?</span></li>
</ul>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Uniqueness:</b><span style="font-weight: 400;"> Are there duplicate records that could skew analysis and operations?</span></li>
</ul>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Validity:</b><span style="font-weight: 400;"> Does the data conform to the defined format, type, and range (e.g., a valid email address format)?</span></li>
</ul>
<p><span style="font-weight: 400;">Assessing and improving data across these dimensions is the first step toward building a trusted data ecosystem.</span></p>
<h3><span style="font-weight: 400;">Defining Data Governance: The strategic framework for </span><span style="font-weight: 400;">c</span><span style="font-weight: 400;">ontrol and value</span></h3>
<p><b>Data governance frameworks</b><span style="font-weight: 400;"> provide the structure for managing an organization&#8217;s data assets. This is not about restricting access but about enabling responsible use. A comprehensive framework establishes the necessary policies, standards, procedures, and controls. It clearly defines who can take what action, with which data, under what circumstances, and using which methods. These </span><b>Data policies</b><span style="font-weight: 400;"> are the rulebook that guides every user in the organization, ensuring that data is handled securely, ethically, and in a way that maximizes its value while minimizing risk.</span></p>
<p><a href="https://inferenz.ai/blogs/the-far-reaching-impact-of-model-drift-and-its-data-drama/"><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12289" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2025/12/CTA-1.gif" alt="" width="1400" height="378" /></a></p>
<h3 style="font-weight: 400;">The Intertwined Nature: How robust governance ensures data Integrity and quality</h3>
<p><span style="font-weight: 400;">Data governance is the engine that drives </span><b>Data Quality</b><span style="font-weight: 400;">. Without a governance framework, efforts to clean up data are temporary fixes at best. Governance establishes the roles and processes needed to maintain data excellence over time. It defines </span><b>Data stewards</b><span style="font-weight: 400;"> who are accountable for specific data domains, implements procedures for data entry and validation, and provides a mechanism for resolving data issues. This structured approach is what ensures </span><b>Data Integrity</b><span style="font-weight: 400;">: the overall accuracy, consistency, and reliability of data throughout its lifecycle. Governance transforms data quality from a reactive, project-based activity into a proactive, embedded discipline.</span></p>
<p><span style="font-weight: 400;">The Cost of Neglect: Addressing Data Trust Issues and Mitigating Reputational Damage</span></p>
<p><span style="font-weight: 400;">Ignoring data quality and governance carries a steep price. Inaccurate customer data leads to poor service and lost sales. Flawed financial data can result in compliance failures and hefty fines. </span></p>
<p><span style="font-weight: 400;">According to Gartner, the average organization loses </span><a href="https://www.gartner.com/smarterwithgartner/how-to-create-a-business-case-for-data-quality-improvement"><span style="font-weight: 400;">$12.9 million annually</span></a><span style="font-weight: 400;"> due to poor data quality. Operationally, bad data creates immense inefficiency as employees spend valuable time hunting for reliable information or correcting errors. Perhaps most damaging is the erosion of trust. When customers lose faith in your ability to manage their information, or when executives can no longer rely on reports to guide the business, the resulting </span><b>reputational damage</b><span style="font-weight: 400;"> can be irreversible.</span></p>
<h2><span style="font-weight: 400;">Crafting Your Strategic Blueprint: Core Pillars of Effective Governance</span></h2>
<p><span style="font-weight: 400;">An effective data governance program is not a one-size-fits-all solution. It must be a carefully designed blueprint tailored to the organization&#8217;s specific needs, maturity, and strategic goals. However, several core pillars are universally essential for success.</span></p>
<h3><span style="font-weight: 400;">        </span><span style="font-weight: 400;">1. Roles and Responsibilities: Empowering Data Stewardship and Leadership</span></h3>
<p><span style="font-weight: 400;">Data governance is a team sport that requires clear accountability. A successful program establishes a hierarchy of roles, starting with executive sponsorship from a </span><b>Chief Data Officer (CDO)</b><span style="font-weight: 400;"> or a similar leader who champions the vision. The most critical on-the-ground role is that of </span><b>Data stewards</b><span style="font-weight: 400;">. These individuals, typically business experts from various departments, are entrusted with overseeing specific organizational data assets. They are responsible for defining data standards, monitoring quality, and ensuring that </span><b>Data policies</b><span style="font-weight: 400;"> are followed within their domain, acting as the crucial link between IT and the business.</span></p>
<h3><span style="font-weight: 400;">        2. Master Data Management (MDM): Achieving a Single, Trusted View of Key Data</span></h3>
<p><span style="font-weight: 400;">Many organizations struggle with fragmented data, where information about a single customer, product, or supplier exists in multiple, often conflicting, versions across different systems. </span><b>Master data management</b><span style="font-weight: 400;"> (MDM) is the discipline and technology used to resolve this chaos. MDM creates a single, authoritative &#8220;golden record&#8221; for critical data entities. </span><b>By creating a central, trusted source of master data, organizations remove inconsistencies. They simplify processes. They make sure all analytics and decisions are based on a shared, accurate view of the business.</b></p>
<h3>        3. <span style="font-weight: 400;">Designing Your Target Operating Model for Data Governance: Structure and Workflow</span></h3>
<p><span style="font-weight: 400;">A Target Operating Model (TOM) for data governance outlines how people, processes, and technology will work together to execute the governance strategy. It defines the structure of the governance council or committee, the workflows for data issue resolution, and the processes for creating and enforcing policies. The TOM serves as the practical implementation plan, detailing how governance will be embedded into the daily operations of the business. It clarifies reporting lines, meeting cadences, and the escalation paths for data-related issues, turning abstract policy into concrete action.</span></p>
<h3>        4. <span style="font-weight: 400;">The Data Lifecycle: Ensuring Quality and Governance from Inception to Archival</span></h3>
<p><span style="font-weight: 400;">Data is not static; it has a lifecycle that begins with its creation and ends with its eventual archival or deletion. Applying data quality and governance principles consistently across this entire journey is essential for maintaining trust and value over time.</span></p>
<h2><span style="font-weight: 400;">Holistic Data Lifecycle Management: A Continuous Journey</span></h2>
<p><span style="font-weight: 400;">Effective </span><b>data lifecycle management</b><span style="font-weight: 400;"> requires a holistic view. This includes managing data creation, storage, usage, sharing, and eventual retirement. Governance procedures must be applied at each stage. For example, data quality checks should be implemented at the point of data entry, access controls must govern its use, and retention policies should dictate how long it is stored. This continuous oversight ensures that </span><b>Data Integrity</b><span style="font-weight: 400;"> is maintained from start to finish.</span></p>
<h2><span style="font-weight: 400;">Data Lineage: Tracing Data&#8217;s Journey and Transformations</span></h2>
<p><b>Data lineage</b><span style="font-weight: 400;"> provides a complete audit trail of data&#8217;s journey through an organization&#8217;s systems. It documents where data originated, what transformations it underwent, and how it is used in various reports and applications. This visibility is crucial for building trust. </span><b>Data lineage is essential for fixing errors. It helps analyze the impact before system changes. It also meets rules for tracking data</b><span style="font-weight: 400;"> for </span><b>regulatory compliance</b><span style="font-weight: 400;">. When a user can see the source and history of a data point, they have more confidence in its accuracy.</span></p>
<h2><span style="font-weight: 400;">Quality and Governance in Modern Data Architectures</span></h2>
<p><span style="font-weight: 400;">The rise of </span><b>big data</b><span style="font-weight: 400;"> technologies, </span><b>Data lakes</b><span style="font-weight: 400;">, and </span><b>Cloud computing</b><span style="font-weight: 400;"> has introduced new challenges for governance. The sheer volume, velocity, and variety of data make manual oversight impossible. To adapt, modern governance frameworks must </span><b>use metadata management tools to automatically list data assets in a data lake. Implement governance controls within cloud platforms. Design a &#8220;data middle platform&#8221; that enforces policies and quality checks on data as it moves between systems.</b><span style="font-weight: 400;"> This ensures a single, governed </span><b>Data Lake</b><span style="font-weight: 400;"> environment rather than a data swamp.</span></p>
<h3><span style="font-weight: 400;">Managing Data Migration and Integration with Quality in Mind</span></h3>
<p><span style="font-weight: 400;">Data migration and system integration projects are high-risk moments for Data Quality. Moving data between systems without proper planning can introduce errors and corrupt information. A robust governance framework is essential to guide these projects. It requires data profiling before migration to find quality problems. It sets clear mapping rules for integration. It demands thorough checks and reconciliation after moving data. This ensures no data is lost or damaged during transfer.</span></p>
<h3><span style="font-weight: 400;">Driving Business Value: Turning Trustworthy Data into Strategic Advantage</span></h3>
<p><span style="font-weight: 400;">The ultimate goal of data quality and governance is not simply to have clean, well-managed data. It is to leverage that data as a strategic asset to drive tangible business outcomes, create competitive differentiation, and foster sustainable growth.</span></p>
<h3><span style="font-weight: 400;">Powering Better Decision-Making and Business Intelligence</span></h3>
<p><span style="font-weight: 400;">The most direct benefit of a strong data governance program is the improvement in strategic and operational decision-making. When executives and managers trust the data in their </span><b>Business intelligence</b><span style="font-weight: 400;"> dashboards and reports, they can make faster, more confident choices. Governed data eliminates the ambiguity and debate over whose numbers are correct, allowing teams to focus on analyzing insights and taking action rather than questioning data validity.</span></p>
<h3><span style="font-weight: 400;">Fueling Advanced Analytics and AI Initiatives</span></h3>
<p><span style="font-weight: 400;">High-quality, well-documented, and easily accessible data is the essential fuel for advanced analytics and </span><b>AI Initiatives</b><span style="font-weight: 400;">. Predictive maintenance models, customer churn predictions, and other machine learning algorithms depend on a rich history of reliable data. </span><b>A governance framework makes sure data is available. </b><span style="font-weight: 400;">It ensures data lineage is clear. It also confirms data is suitable for advanced applications. This greatly raises the chance of success for an organization&#8217;s top projects.</span></p>
<h3><span style="font-weight: 400;">Enhancing Customer and User Experience with Reliable Data</span></h3>
<p><span style="font-weight: 400;">Reliable data is the foundation of a superior customer experience. A single, accurate view of the customer, enabled by MDM, allows for true personalization, targeted marketing, and seamless service interactions. When a user contacts support, they expect the agent to have their complete and correct history. Inaccurate or incomplete data leads to frustrating, disjointed experiences that damage customer loyalty and brand perception.</span></p>
<h3><span style="font-weight: 400;">Optimizing Business Processes and Operational Efficiency</span></h3>
<p><span style="font-weight: 400;">Clean, consistent, and timely data is a powerful catalyst for operational excellence. It streamlines business processes by removing the friction caused by data errors. For example, accurate product data reduces shipping errors in logistics, correct supplier data ensures timely payments in procurement, and valid employee data simplifies HR and payroll processes. These efficiencies compound across the organization, reducing operational costs and freeing up employee time for more value-added activities.</span></p>
<h3><span style="font-weight: 400;">Enabling Data Accessibility and Responsible Data Sharing</span></h3>
<p><span style="font-weight: 400;">A common misconception is that governance is about locking data down. </span><b>In reality, good governance supports responsible data access.</b><span style="font-weight: 400;"> By establishing clear ownership, security classifications, and access policies, governance creates a framework for </span><b>Data Accessibility</b><span style="font-weight: 400;"> where data can be shared confidently and securely across the organization. This &#8220;data democratization&#8221; empowers more users to access the data they need to perform their jobs effectively while ensuring that sensitive information is protected.</span></p>
<h3><span style="font-weight: 400;">Mitigating Risk &amp; Ensuring Trust: The Compliance and Security Imperative</span></h3>
<p><span style="font-weight: 400;">In an increasingly regulated world, robust data governance is no longer optional; it is a fundamental component of risk management. It provides the necessary controls and oversight to protect the organization from regulatory penalties, security breaches, and the associated reputational fallout.</span></p>
<h3><span style="font-weight: 400;">Navigating the Complex Landscape of Regulatory Compliance</span></h3>
<p><span style="font-weight: 400;">Organizations today face a complex web of </span><b>privacy laws</b><span style="font-weight: 400;"> and data protection regulations, such as the EU&#8217;s GDPR and the California Consumer Privacy Act (CCPA). Adhering to these rules requires a deep understanding of what data is collected, where it is stored, and how it is used. </span><b>Data governance </b><span style="font-weight: 400;">frameworks manage regulatory compliance. They document data processing activities, handle consent, and enforce policies. These ensure data is used according to legal rules.</span></p>
<h3><span style="font-weight: 400;">Proactive Risk Management: Data Audit and Data Observability for Continuous Oversight</span></h3>
<p><span style="font-weight: 400;">Instead of reacting to data breaches or quality failures, leading organizations are adopting proactive risk management strategies. This includes regular data audits to assess compliance with internal policies and external regulations. The emerging field of </span><b>Data Observability</b><span style="font-weight: 400;"> goes a step further, using automated tools to continuously monitor the health of data pipelines and systems. This provides real-time alerts on data quality degradation, schema changes, or anomalous data patterns, allowing teams to identify and resolve issues before they impact the business.</span></p>
<h3><span style="font-weight: 400;">Establishing Clear Data Issue Escalation and Resolution Processes</span></h3>
<p><span style="font-weight: 400;">Even with the best controls, data issues will inevitably arise. A key function of data governance is to establish clear, efficient procedures for identifying, escalating, and resolving these issues. A defined data issue escalation path ensures that when a user spots a problem, they know exactly who to report it to. This process guarantees that the right </span><b>Data stewards</b><span style="font-weight: 400;"> and technical teams are engaged quickly to perform root cause analysis and implement a lasting solution, preventing the same issue from recurring.</span></p>
<h2><span style="font-weight: 400;">The Human Element &amp; Cultural Transformation: Building a Data-Driven Organization</span></h2>
<p><span style="font-weight: 400;">Ultimately, technology and policies are only part of the solution. Achieving a truly data-driven organization requires a cultural transformation. It means fostering a shared sense of responsibility for data quality across all departments and empowering every employee with the skills and knowledge to treat data as a critical enterprise asset. This cultural shift, supported by strong leadership and continuous training, is what turns a governance blueprint into a living, breathing reality.</span></p>
<p><a href="https://inferenz.ai/contact-us/" target="_blank" rel="noopener"><img loading="lazy" decoding="async" class="alignleft wp-image-12290 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2025/12/CTA-2.gif" alt="" width="1400" height="378" /></a></p>
<h2><span style="font-weight: 400;">Conclusion</span></h2>
<p><span style="font-weight: 400;">Data quality and governance are not mere technical exercises or compliance hurdles; they are the strategic blueprint for sustainable success in the digital age. By implementing a robust framework built on clear roles, effective processes, and enabling technologies like </span><b>Master data management</b><span style="font-weight: 400;">, organizations can transform their data from a chaotic liability into their most powerful asset. </span><b>This change helps make smarter decisions. It improves the customer experience and increases operational efficiency. It also creates a necessary base for successful AI initiatives.</b></p>
<p><span style="font-weight: 400;">The journey begins by treating data as a core business function, not an IT afterthought. It requires building a culture of accountability where everyone understands their role in preserving </span><b>Data Integrity</b><span style="font-weight: 400;"> and upholding quality. By committing to this blueprint, organizations can confidently navigate the complexities of the modern data landscape, mitigate risk, and unlock the full potential of their information assets. </span><b>By investing in this plan, your organization can do more than manage data. It can actively use data to find new ways to innovate. It can reduce risks and gain a lasting competitive edge.</b></p>
<p>The post <a href="https://inferenz.ai/blogs/data-quality-and-governance-for-scalable-and-sustainable-growth/">Data Quality &#038; Governance: The Strategic Blueprint for Sustainable Organizational Success</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>The Far-Reaching Impact of Model Drift and its Data Drama</title>
		<link>https://inferenz.ai/blogs/the-far-reaching-impact-of-model-drift-and-its-data-drama/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Mon, 17 Nov 2025 08:57:00 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://inferenz.ai/?p=12226</guid>

					<description><![CDATA[<p>Model drift is more than a real data science headache, it’s a silent business killer. When the data your AI relies on changes, predictions falter, decisions suffer, and trust erodes.</p>
<p>The post <a href="https://inferenz.ai/blogs/the-far-reaching-impact-of-model-drift-and-its-data-drama/">The Far-Reaching Impact of Model Drift and its Data Drama</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p></p>
<h2><span class="TextRun SCXW37838156 BCX0" lang="EN-IN" xml:lang="EN-IN" data-contrast="auto"><span class="NormalTextRun SCXW37838156 BCX0">Background Summary</span></span></h2>
<p><span style="font-weight: 400;">Model drift is more than a real data science headache, it’s a silent business killer. When the data your AI relies on changes, predictions falter, decisions suffer, and trust erodes. This guide explains what drift is, why it affects every industry, and how a mix of smart monitoring, robust data pipelines, and AI-powered cleaning tools can keep your models performing at their peak.</span>-<span style="font-weight: 400;">Imagine launching a new product, rolling out a service upgrade, or opening a flagship store after months of preparation, only to find customer complaints piling up because something invisible changed behind the scenes. In AI, that invisible culprit is often model drift.</span></p>
<p><span style="font-weight: 400;">Your model worked perfectly in testing. Predictions were accurate, dashboards lit up with promising KPIs.  But months later, results dip, costs climb, and customer trust erodes. What changed? </span></p>
<p><span style="font-weight: 400;">The data feeding your model no longer reflects the real world it serves. </span></p>
<p><span style="font-weight: 400;">This article breaks down why that happens, why it matters to every industry, and how modern tools can stop drift before it damages outcomes.</span></p>
<h2><span style="font-weight: 400;">What is “Data Drama”?</span></h2>
<p><span style="font-weight: 400;">“Data drama” means wrestling with disorganized, inconsistent, or incomplete data when building AI solutions, leading to model drift. Model drift refers to </span><b>the degradation of a model’s performance over time due to changes in data distribution or the environment</b><span style="font-weight: 400;"> it operates in.</span></p>
<p><span style="font-weight: 400;">Think of it as junk in the trunk: if your AI is the car, bad data makes for a bumpy ride, no matter how powerful the engine is.</span></p>
<p><span style="font-weight: 400;">Picture a hospital that wants to use AI to predict patient health risks:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Patient names are sometimes written “Jon Smith,” “John Smith,” or “J. Smith.”</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Some records are missing phone numbers or have outdated addresses.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">The hospital’s old records are stored in paper files or weird formats.</span></li>
</ul>
<p><span style="font-weight: 400;">Even if the AI is “smart,” it struggles to learn from such confusing information. There are three primary types of drifts that affect the scenarios:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Data drift (covariate shift):</b><span style="font-weight: 400;"> The input distribution P(x) changes. Example: new user behavior, seasonal trends, new data sources.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Concept drift:</b><span style="font-weight: 400;"> The relationship between features and target P(y</span><span style="font-weight: 400;">∣</span><span style="font-weight: 400;">x) changes. Example: fraud tactics evolve customer churn reasons shift.</span></li>
</ul>
<p><b>Label drift (prior probability shift):</b><span style="font-weight: 400;"> The distribution of P(y) changes. Common in imbalanced classification tasks.</span></p>
<h2><span style="font-weight: 400;">Why is this a problem?</span></h2>
<p><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12234" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2025/11/Why-is-this-a-problem.jpg" alt="" width="1440" height="1029" /></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Silent failures:</b><span style="font-weight: 400;"> Drift isn’t always obvious models can keep running, just poorly.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Bad decisions:</b><span style="font-weight: 400;"> In finance, healthcare, or logistics, this can mean misdiagnoses, delays, or big financial losses.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Customer frustration:</b><span style="font-weight: 400;"> Imagine getting your credit card blocked for every vacation you take.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Wasted resources:</b><span style="font-weight: 400;"> Fixing a broken model after damage is harder (and costlier) than preventing it.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Time wasted:</b><span style="font-weight: 400;"> Engineers spend up to 80% of their time cleaning data instead of building useful solutions.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Hidden mistakes: </b><span style="font-weight: 400;">Flawed data can make the AI give wrong answers—like approving the wrong credit card application or missing a fraud alert.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Loss of trust:</b> If the AI presents inaccurate results, users quickly lose faith in the technology.</li>
</ul>
<h3><span style="font-weight: 400;">Why is it hard to catch?</span></h3>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Most production pipelines don’t monitor </span><b>live feature distributions</b><span style="font-weight: 400;"> or </span><b>prediction confidence</b><span style="font-weight: 400;">.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Business KPIs may degrade before engineers notice any statistical performance drop.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Retraining isn’t always feasible daily, especially without label feedback loops.</span></li>
</ul>
<h2><span style="font-weight: 400;">How can we solve the data drama?</span></h2>
<p><span style="font-weight: 400;">Today, AI itself helps clean and fix messy data, making life easier for both techies and non-techies. Here’s a step-by-step technical approach for managing drift in production systems: </span></p>
<h3><span style="font-weight: 400;">           1. Track key statistical metrics on input data:</span></h3>
<ul>
<li style="list-style-type: none;">
<ul>
<li style="list-style-type: none;">
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Population stability index (PSI)</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Kullback-leibler divergence (KL Divergence)</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Kolmogorov-smirnov (KS) test</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Wasserstein distance (for continuous features)</span></li>
</ul>
</li>
</ul>
</li>
</ul>
<p><b>Implementation example:</b></p>
<p><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12229" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2025/11/Code1.jpg" alt="" width="1440" height="525" /></p>
<p><b>Tools</b><span style="font-weight: 400;">: Evidently AI, WhyLabs, Arize AI</span></p>
<h3><span style="font-weight: 400;">          2. Monitoring model performance without labels</span></h3>
<p><span style="font-weight: 400;">If you can’t get real-time labels, use </span><b>proxy indicators</b><span style="font-weight: 400;">:</span></p>
<ol>
<li style="list-style-type: none;">
<ol>
<li style="list-style-type: none;">
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Confidence score distributions</b><span style="font-weight: 400;"> (are they shifting?)</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Prediction entropy</b><span style="font-weight: 400;"> or </span><b>uncertainty variance</b></li>
<li style="font-weight: 400;" aria-level="1"><b>Output class distribution shift</b></li>
</ul>
</li>
</ol>
</li>
</ol>
<p><b>Example using fiddler AI</b><span style="font-weight: 400;">:</span></p>
<p><span style="font-weight: 400;"><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12230" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2025/11/Code2.jpg" alt="" width="1440" height="404" /><br />
# Detect divergence from training output distributions</span><span style="font-weight: 400;"><br />
</span></p>
<h3><span style="font-weight: 400;">          3. Retraining pipelines &amp; model registry integration</span></h3>
<p><span style="font-weight: 400;">Build retraining workflows that:</span></p>
<ol>
<li style="list-style-type: none;">
<ol>
<li style="list-style-type: none;">
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Pull recent production data</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Recompute features</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Revalidate on held-out test sets</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Re-register the model with metadata</span></li>
</ul>
</li>
</ol>
</li>
</ol>
<p><b>Example stack:</b></p>
<ol>
<li style="list-style-type: none;">
<ol>
<li style="list-style-type: none;">
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Feature store</b><span style="font-weight: 400;">: Feast / Tecton</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Training pipelines</b><span style="font-weight: 400;">: MLflow / SageMaker Pipelines / Vertex AI</span></li>
<li style="font-weight: 400;" aria-level="1"><b>CI/CD</b><span style="font-weight: 400;">: GitHub Actions + DVC</span></li>
</ul>
</li>
</ol>
</li>
</ol>
<p><b>Registry</b><span style="font-weight: 400;">: MLflow or SageMaker Model Registry</span><br />
<a href="https://inferenz.ai/blogs/data-observability-in-snowflake-a-hands-on-technical-guide/"><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12231" style="width: 100%; margin-bottom: 25px;" src="https://inferenz.ai/wp-content/uploads/2025/11/CTA-1-1.gif" alt="" width="1400" height="378" /></a></p>
<h2><span style="font-weight: 400;">Tools &amp; solutions </span></h2>
<p><span style="font-weight: 400;">This is broken down by stages of the solution pipeline:</span></p>
<h3>1.<span style="font-weight: 400;">Understanding what data is missing</span></h3>
<p><span style="font-weight: 400;">Before solving the problem, you need to </span><b>identify what is missing</b><span style="font-weight: 400;"> or </span><b>irrelevant</b><span style="font-weight: 400;"> in your dataset.</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Tool</b></td>
<td><b>Purpose</b></td>
<td><b>Features</b></td>
</tr>
<tr>
<td><b>Great expectations</b></td>
<td><span style="font-weight: 400;">Data profiling, testing, validation</span></td>
<td><span style="font-weight: 400;">Detects missing values, schema mismatches, unexpected distributions</span></td>
</tr>
<tr>
<td><b>Pandas profiling / YData profiling</b></td>
<td><span style="font-weight: 400;">Exploratory data analysis</span></td>
<td><span style="font-weight: 400;">Generates auto-EDA reports; useful to check data completeness</span></td>
</tr>
<tr>
<td><b>Data contracts (openLineage, dataplex)</b></td>
<td><span style="font-weight: 400;">Define expected data schema and sources</span></td>
<td><span style="font-weight: 400;">Ensures the data you need is being collected consistently</span></td>
</tr>
</tbody>
</table>
</div>
<p>&nbsp;</p>
<h3><span style="font-weight: 400;"> 2. Data collection &amp; logging infrastructure</span></h3>
<p><span style="font-weight: 400;">To fix missing data, you need to </span><b>collect more meaningful, raw, or contextual signals</b><span style="font-weight: 400;">—especially behavioral or operational data.</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Tool</b></td>
<td><b>Use Case</b></td>
<td><b>Integration</b></td>
</tr>
<tr>
<td><b>Apache kafka</b></td>
<td><span style="font-weight: 400;">Real-time event logging</span></td>
<td><span style="font-weight: 400;">Captures user behavior, app events, support logs</span></td>
</tr>
<tr>
<td><b>Snowplow analytics</b></td>
<td><span style="font-weight: 400;">User tracking infrastructure</span></td>
<td><span style="font-weight: 400;">Web/mobile event tracking pipeline for custom behaviors</span></td>
</tr>
<tr>
<td><b>Segment</b></td>
<td><span style="font-weight: 400;">Customer data platform</span></td>
<td><span style="font-weight: 400;">Collects customer touchpoints and routes to data warehouses</span></td>
</tr>
<tr>
<td><b>OpenTelemetry</b></td>
<td><span style="font-weight: 400;">Observability for services</span></td>
<td><span style="font-weight: 400;">Track service logs, latency, API calls tied to user sessions</span></td>
</tr>
<tr>
<td><b>Fluentd / Logstash</b></td>
<td><span style="font-weight: 400;">Log collectors</span></td>
<td><span style="font-weight: 400;">Integrate service and system logs into pipelines for ML use</span></td>
</tr>
</tbody>
</table>
</div>
<p>&nbsp;</p>
<h3><span style="font-weight: 400;">3. Feature engineering &amp; enrichment</span></h3>
<p><span style="font-weight: 400;">Once the relevant data is collected, you’ll need to </span><b>transform it into usable features</b><span style="font-weight: 400;">—especially across systems.</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Tool</b></td>
<td><b>Use Case</b></td>
<td><b>Notes</b></td>
</tr>
<tr>
<td><b>Feast</b></td>
<td><span style="font-weight: 400;">Open-source feature store</span></td>
<td><span style="font-weight: 400;">Manages real-time and offline features, auto-syncs with models</span></td>
</tr>
<tr>
<td><b>Tecton</b></td>
<td><span style="font-weight: 400;">Enterprise-grade feature platform</span></td>
<td><span style="font-weight: 400;">Centralized feature pipelines, freshness tracking, time-travel</span></td>
</tr>
<tr>
<td><b>Databricks feature store</b></td>
<td><span style="font-weight: 400;">Native with Delta Lake</span></td>
<td><span style="font-weight: 400;">Integrates with MLflow, auto-tracks lineage</span></td>
</tr>
<tr>
<td><b>DBT + Snowflake</b></td>
<td><span style="font-weight: 400;">Feature pipelines via SQL</span></td>
<td><span style="font-weight: 400;">Great for tabular/business data pipelines</span></td>
</tr>
<tr>
<td><b>Google vertex AI feature store</b></td>
<td><span style="font-weight: 400;">Fully managed</span></td>
<td><span style="font-weight: 400;">Ideal for GCP users with built-in monitoring</span></td>
</tr>
</tbody>
</table>
</div>
<p>&nbsp;</p>
<h3><span style="font-weight: 400;">4. External &amp; third-party data integration</span></h3>
<p><span style="font-weight: 400;">Some of the </span><b>most relevant data may come from external APIs or third-party sources</b><span style="font-weight: 400;">, especially in domains like finance, health, logistics, and retail.</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Data type</b></td>
<td><b>Tools / APIs</b></td>
</tr>
<tr>
<td><b>Weather, location</b></td>
<td><span style="font-weight: 400;">OpenWeatherMap, HERE Maps, NOAA APIs</span></td>
</tr>
<tr>
<td><b>Financial scores</b></td>
<td><span style="font-weight: 400;">Experian, Equifax APIs</span></td>
</tr>
<tr>
<td><b>News/sentiment</b></td>
<td><span style="font-weight: 400;">GDELT, Google Trends, LexisNexis</span></td>
</tr>
<tr>
<td><b>Support tickets</b></td>
<td><span style="font-weight: 400;">Zendesk API, Intercom API</span></td>
</tr>
<tr>
<td><b>Social/feedback</b></td>
<td><span style="font-weight: 400;">Trustpilot API, Twitter API, App Store reviews</span></td>
</tr>
</tbody>
</table>
</div>
<p>&nbsp;</p>
<h3><span style="font-weight: 400;">5. Data observability &amp; monitoring</span></h3>
<p><span style="font-weight: 400;">Once new data is flowing, ensure its </span><b>quality, freshness, and availability</b><span style="font-weight: 400;"> remain intact.</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Tool</b></td>
<td><b>Capabilities</b></td>
</tr>
<tr>
<td><b>Evidently AI</b></td>
<td><span style="font-weight: 400;">Data drift, feature distribution, missing value alerts</span></td>
</tr>
<tr>
<td><b>WhyLabs</b></td>
<td><span style="font-weight: 400;">Real-time observability for structured + unstructured data</span></td>
</tr>
<tr>
<td><b>Monte Carlo</b></td>
<td><span style="font-weight: 400;">Data lineage, freshness monitoring across pipelines</span></td>
</tr>
<tr>
<td><b>Soda.io</b></td>
<td><span style="font-weight: 400;">Data quality monitoring with alerts and testing</span></td>
</tr>
<tr>
<td><b>Datafold</b></td>
<td><span style="font-weight: 400;">Data diffing and schema change tracking</span></td>
</tr>
</tbody>
</table>
</div>
<p>&nbsp;</p>
<h3><span style="font-weight: 400;">6. Explainability &amp; impact analysis</span></h3>
<p><span style="font-weight: 400;">You want to make sure </span><b>your added features are actually helping</b><span style="font-weight: 400;"> the model and understand their impact.</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Tool</b></td>
<td><b>Use Case</b></td>
</tr>
<tr>
<td><b>SHAP / LIME</b></td>
<td><span style="font-weight: 400;">Explain model decisions feature-wise</span></td>
</tr>
<tr>
<td><b>Fiddler AI</b></td>
<td><span style="font-weight: 400;">Combines drift detection + explainability</span></td>
</tr>
<tr>
<td><b>Arize AI</b></td>
<td><span style="font-weight: 400;">Real-time monitoring and root-cause drift analysis</span></td>
</tr>
<tr>
<td><b>Captum (for PyTorch)</b></td>
<td><span style="font-weight: 400;">Deep learning explainability library</span></td>
</tr>
</tbody>
</table>
</div>
<p>&nbsp;</p>
<h2><span style="font-weight: 400;">Why model drift is every business’s problem</span></h2>
<p><span style="font-weight: 400;">Model drift may sound like a technical glitch, but its consequences ripple across industries in ways that hurt revenue, efficiency, and trust.</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Healthcare</b><span style="font-weight: 400;"> – A drifted model can misread patient risk levels, causing </span><i><span style="font-weight: 400;">missed diagnoses</span></i><span style="font-weight: 400;">, delayed interventions, or unnecessary tests. In critical care, this can directly affect patient outcomes.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Finance</b><span style="font-weight: 400;"> – Inconsistent data patterns can produce </span><i><span style="font-weight: 400;">incorrect credit scoring</span></i><span style="font-weight: 400;"> or flag legitimate transactions as fraudulent, frustrating customers and damaging loyalty.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Retail &amp; E-commerce</b><span style="font-weight: 400;"> – Changing buying behavior or seasonal demand shifts can lead to </span><i><span style="font-weight: 400;">inaccurate demand forecasts</span></i><span style="font-weight: 400;">, resulting in overstock that ties up cash or stockouts that push customers to competitors.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Manufacturing &amp; supply chain</b><span style="font-weight: 400;"> – Predictive maintenance models can miss early signs of equipment wear, leading to </span><i><span style="font-weight: 400;">unplanned downtime</span></i><span style="font-weight: 400;"> that halts production lines.</span></li>
</ul>
<h4><i><span style="font-weight: 400;">The common thread?</span></i></h4>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Revenue impact</b><span style="font-weight: 400;"> – Poor predictions lead to lost sales opportunities and operational waste.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Compliance risk</b><span style="font-weight: 400;"> – In regulated sectors, drift can create breaches in reporting accuracy or fairness obligations.</span></li>
</ul>
<p><b>Brand reputation</b><span style="font-weight: 400;"> – Customers and partners lose trust if decisions feel inconsistent or incorrect.</span></p>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignleft wp-image-12232 size-full" style="width: 100%; margin-bottom: 25px;" src="https://inferenz.ai/wp-content/uploads/2025/11/CTA-2-1.gif" alt="" width="1400" height="378" /></a></p>
<h2><span style="font-weight: 400; margin-top: 30px;">The cost of ignoring model drift</span></h2>
<p><span style="font-weight: 400;">The business case for tackling drift is backed by hard numbers:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><a href="https://www.gartner.com/en/data-analytics/topics/data-quality" target="_blank" rel="noopener"><span style="font-weight: 400;">Data quality issues</span></a><span style="font-weight: 400;"> cost organizations an average of $12.9 million annually.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">For predictive systems, </span><a href="https://iot-analytics.com/predictive-maintenance-market" target="_blank" rel="noopener"><b>downtime</b></a><b> can cost $125,000 per hour</b><span style="font-weight: 400;"> on an average depending on the industry.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Recovery from a drifted model, retraining, redeployment, and regaining lost customer trust, can take </span><b>weeks to months</b><span style="font-weight: 400;">, costing far more than prevention.</span></li>
</ul>
<p><span style="font-weight: 400;">Implementing automated drift detection can reduce model troubleshooting time drastically.  Early intervention can prevent revenue losses in industries where decisions are AI-driven.</span></p>
<p><span style="font-weight: 400;">In other words, the cost of </span><i><span style="font-weight: 400;">not</span></i><span style="font-weight: 400;"> acting is often several times higher than the cost of building proactive safeguards.</span></p>
<h2><span style="font-weight: 400;">From detection to prevention</span></h2>
<p><span style="font-weight: 400;">Drift management is about more than catching problems, it’s about designing systems that keep models healthy and relevant from the start.</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Approach</b></td>
<td><b>What It Looks Like</b></td>
<td><b>Outcome</b></td>
</tr>
<tr>
<td><b>Reactive</b></td>
<td><span style="font-weight: 400;">Model performance dips → business KPIs drop → engineers scramble to investigate.</span></td>
<td><span style="font-weight: 400;">Higher downtime, lost revenue, longer recovery cycles.</span></td>
</tr>
<tr>
<td><b>Proactive</b></td>
<td><span style="font-weight: 400;">Continuous monitoring of data and predictions → alerts trigger retraining before business impact.</span></td>
<td><span style="font-weight: 400;">Minimal disruption, sustained model accuracy, preserved customer trust.</span></td>
</tr>
</tbody>
</table>
</div>
<p><b>Why proactive wins:</b></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Reduces firefighting and emergency fixes.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Ensures AI systems adapt alongside market or operational changes.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Turns drift management into a </span><b>competitive advantage</b><span style="font-weight: 400;">, keeping predictions accurate while competitors struggle with outdated models.</span></li>
</ul>
<p>&nbsp;</p>
<h2><span style="font-weight: 400;">Takeaway</span></h2>
<p><span style="font-weight: 400;">In fast-moving markets, your AI is only as good as the data it learns from. Drift happens quietly, but its effects ripple loudly across customer experiences, operational efficiency, and revenue. By combining continuous monitoring with adaptive retraining, businesses can turn model drift from a costly disruption into a controlled, measurable process.</span></p>
<p><span style="font-weight: 400;">The real win is beyond the fact that it fixes broken predictions. Now you can build AI systems that grow alongside your business, staying relevant and reliable in any market condition.</span></p>
<p>The post <a href="https://inferenz.ai/blogs/the-far-reaching-impact-of-model-drift-and-its-data-drama/">The Far-Reaching Impact of Model Drift and its Data Drama</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>QA in the Modern Data Stack: Using Python, Zephyr Scale &#038; Unity Catalog for End-to-End Quality Assurance</title>
		<link>https://inferenz.ai/blogs/qa-in-the-modern-data-stack-using-python-zephyr-scale-unity-catalog-for-end-to-end-quality-assurance/</link>
		
		<dc:creator><![CDATA[inferenz.manage]]></dc:creator>
		<pubDate>Wed, 29 Oct 2025 04:48:57 +0000</pubDate>
				<category><![CDATA[Data & Cloud Migration]]></category>
		<guid isPermaLink="false">https://inferenz.ai/?p=12048</guid>

					<description><![CDATA[<p>&#160; Integrated QA framework using Python, Zephyr Scale &#38; Unity Catalog Introduction Quality Assurance (QA) in the software world has moved beyond functional testing and interface validation. As modern enterprises shift toward data-centric architectures and cloud-native platforms, QA now involves ensuring data accuracy, integrity, governance, and system compliance end to end. In a recent enterprise [&#8230;]</p>
<p>The post <a href="https://inferenz.ai/blogs/qa-in-the-modern-data-stack-using-python-zephyr-scale-unity-catalog-for-end-to-end-quality-assurance/">QA in the Modern Data Stack: Using Python, Zephyr Scale &#038; Unity Catalog for End-to-End Quality Assurance</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>&nbsp;</p>
<p><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12052" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2025/10/Integrated-QA-Framework-Using-Python-Zephyr-Scale-Unity-Catalog.jpg" alt="" width="1440" height="1029" /><em>Integrated QA framework using Python, Zephyr Scale &amp; Unity Catalog</em></p>
<h2><span style="font-weight: 400;">Introduction</span></h2>
<p><span style="font-weight: 400;">Quality Assurance (QA) in the software world has moved beyond functional testing and interface validation. As modern enterprises shift toward </span><b>data-centric architectures and cloud-native platforms</b><span style="font-weight: 400;">, QA now involves ensuring </span><b>data accuracy, integrity, governance, and system compliance</b><span style="font-weight: 400;"> end to end.</span></p>
<p><span style="font-weight: 400;">In a recent enterprise project, I worked on migrating a </span><b>legacy Customer Relationship Management (CRM)</b><span style="font-weight: 400;"> system to </span><b>Microsoft Dynamics 365 (MS D365)</b><span style="font-weight: 400;">. It wasn’t merely a technology shift. It involved moving large data volumes, aligning new business rules, setting up strong governance layers, and ensuring uninterrupted business operations.</span></p>
<p><span style="font-weight: 400;">In this article, I’ll share how QA was handled across this transformation using </span><b>Zephyr Scale</b><span style="font-weight: 400;"> for test management, </span><b>Python</b><span style="font-weight: 400;"> for automation, and </span><a href="https://inferenz.ai/blogs/databricks-unity-catalog-building-a-unified-data-governance-layer-in-modern-data-platforms/"><b>Databricks Unity Catalog</b></a><span style="font-weight: 400;"> for governance and access control.</span></p>
<h2><span style="font-weight: 400;">QA challenges in migrating to Microsoft Dynamics 365</span></h2>
<p><span style="font-weight: 400;">Migrating from a legacy CRM to a modern cloud platform brings unique QA challenges. The main focus areas included:</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b><i>Focus Area</i></b></td>
<td><b>QA Objective</b></td>
<td><b>Common Issues</b></td>
</tr>
<tr>
<td><b><i>Data Validation</i></b></td>
<td><span style="font-weight: 400;">Ensure data integrity and accuracy post-migration</span></td>
<td><span style="font-weight: 400;">Missing, duplicate, or corrupted records</span></td>
</tr>
<tr>
<td><b><i>Functional Testing</i></b></td>
<td><span style="font-weight: 400;">Validate end-to-end workflows across Bronze → Silver → Gold layers</span></td>
<td><span style="font-weight: 400;">Breaks in business logic or incomplete process flow</span></td>
</tr>
<tr>
<td><b><i>Integration Testing</i></b></td>
<td><span style="font-weight: 400;">Verify KPI accuracy in downstream systems</span></td>
<td><span style="font-weight: 400;">Data mismatch or inconsistent calculations</span></td>
</tr>
</tbody>
</table>
</div>
<p><span style="font-weight: 400;">This was my first experience in a </span><b>hybrid QA setup</b><span style="font-weight: 400;">—where data engineering and cloud CRM validation worked together. Automation became essential from the start.</span></p>
<h2><span style="font-weight: 400;">Test management with Zephyr Scale in Jira</span></h2>
<p><span style="font-weight: 400;">We used </span><b>Zephyr Scale</b><span style="font-weight: 400;"> within </span><b>Jira</b><span style="font-weight: 400;"> to manage all QA activities. It ensured complete traceability from </span><i><span style="font-weight: 400;">test case creation → execution → defect resolution.</span></i></p>
<p><span style="font-weight: 400;">The test planning followed an iterative Agile structure:</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Sprint</b></td>
<td><b>Phase</b></td>
<td><b>Description</b></td>
</tr>
<tr>
<td><b>Sprint 1</b></td>
<td><span style="font-weight: 400;">System Integration Testing (SIT)</span></td>
<td><span style="font-weight: 400;">Validation of data flow, transformations, and business rules</span></td>
</tr>
<tr>
<td><b>Sprint 2</b></td>
<td><span style="font-weight: 400;">User Acceptance Testing (UAT)</span></td>
<td><span style="font-weight: 400;">Final stage readiness checks before production deployment</span></td>
</tr>
</tbody>
</table>
</div>
<p><b>Sample migration test case</b></p>
<p><b>Objective:</b><span style="font-weight: 400;"> Validate that data from the Bronze layer is accurately transferred to the Silver layer.</span></p>
<p><b>Steps:</b></p>
<ol>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Query record counts in the Bronze schema.  </span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Query corresponding counts in the Silver schema.  </span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Compare totals and sample values.  </span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Confirm no data loss or duplication.</span></li>
</ol>
<p><span style="font-weight: 400;">Zephyr Scale offered complete visibility—allowing both QA and business teams to align quickly and demonstrate readiness during go-live reviews.</span></p>
<p><a href="https://inferenz.ai/case-studies/"><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12050" style="width: 100; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2025/10/CTA-1-2.gif" alt="" width="1400" height="378" /></a></p>
<h2><span style="font-weight: 400;">Writing effective test scenarios and cases</span></h2>
<p><span style="font-weight: 400;">In a data migration project, QA must cover both systems—the old CRM and the new MS D365—along with the underlying </span><b>Databricks Lakehouse layers.</b></p>
<p><span style="font-weight: 400;">The following scenarios formed the backbone of our testing effort:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><b>Data validation:</b><span style="font-weight: 400;"> Ensuring every record from the old subscription is fully and accurately migrated.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Schema validation:</b><span style="font-weight: 400;"> Confirming the data flow through Bronze → Silver layers, with cleansing and normalization (3NF) applied.</span></li>
<li style="font-weight: 400;" aria-level="1"><b>KPI validation:</b><span style="font-weight: 400;"> Verifying 16 business KPIs for accuracy, completeness, and correct duration (annual or quarterly).</span></li>
<li style="font-weight: 400;" aria-level="1"><b>Governance validation:</b><span style="font-weight: 400;"> Checking access permissions, lineage, and audit logs for compliance.</span></li>
</ul>
<p><span style="font-weight: 400;">This structured approach ensured coverage across the technical and business sides of the migration.</span></p>
<h2><span style="font-weight: 400;">QA automation with Python</span></h2>
<p><span style="font-weight: 400;">Manual validation quickly became impractical with large datasets and frequent syncs. Automation was the only sustainable approach.</span></p>
<p><b>Automated checks included:</b></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Record counts between schemas/tables/columns</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Schema conformity checks in migrated tables</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Data Validation from Bronze to Silver to Gold</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Naming convention checks</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Storage location validations</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">KPI Calculations</span></li>
</ul>
<p><span style="font-weight: 400;">This automation saved countless hours and ensured we caught discrepancies quickly.</span></p>
<p><b><i>Sample script:</i></b></p>
<p><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12054" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2025/10/Sample-Script.jpg" alt="" width="1440" height="525" /></p>
<p>These automated tests reduced QA time, enabled early detection of errors, and ensured reliable validation across migration batches.</p>
<h2><span style="font-weight: 400;">Unity Catalog: Governance in the data pipeline</span></h2>
<p><span style="font-weight: 400;">Data governance was as important as data accuracy in this project. Using </span><a href="https://inferenz.ai/blogs/databricks-unity-catalog-building-a-unified-data-governance-layer-in-modern-data-platforms/"><b>Databricks Unity Catalog</b></a><span style="font-weight: 400;">, we centralized security, access, and lineage validation for all datasets.</span></p>
<p><span style="font-weight: 400;">As part of QA, we validated:</span></p>
<div class="table-responsive">
<table style="margin: 0px;">
<tbody>
<tr>
<td><b>Governance Check</b></td>
<td><b>QA Objective</b></td>
</tr>
<tr>
<td><b>Access Control</b></td>
<td><span style="font-weight: 400;">Ensure only authorized users can view Personally Identifiable Information (PII).</span></td>
</tr>
<tr>
<td><b>Schema Locking</b></td>
<td><span style="font-weight: 400;">Validate that schema versions remain consistent across deployments.</span></td>
</tr>
<tr>
<td><b>Audit Logging</b></td>
<td><span style="font-weight: 400;">Confirm all data access events are recorded and retrievable.</span></td>
</tr>
</tbody>
</table>
</div>
<p><span style="font-weight: 400;">Testing with Unity Catalog reinforced compliance while maintaining transparency across teams.</span></p>
<h2><span style="font-weight: 400;">End-to-end QA workflow in the migration</span></h2>
<p><span style="font-weight: 400;">Each tool contributed to the overall assurance model:</span></p>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Step</b></td>
<td><b>Tool Used</b></td>
<td><b>QA Outcome</b></td>
</tr>
<tr>
<td><b>Test scenario creation</b></td>
<td><span style="font-weight: 400;">Zephyr Scale + Jira</span></td>
<td><span style="font-weight: 400;">Linked to user stories for visibility</span></td>
</tr>
<tr>
<td><b>Data validation</b></td>
<td><span style="font-weight: 400;">Python automation</span></td>
<td><span style="font-weight: 400;">Verified migration accuracy</span></td>
</tr>
<tr>
<td><b>Governance checks</b></td>
<td><span style="font-weight: 400;">Unity Catalog</span></td>
<td><span style="font-weight: 400;">Validated access control and data lineage</span></td>
</tr>
<tr>
<td><b>Reporting</b></td>
<td><span style="font-weight: 400;">Zephyr dashboards</span></td>
<td><span style="font-weight: 400;">Weekly QA progress reports</span></td>
</tr>
</tbody>
</table>
<p>&nbsp;</p>
</div>
<p><a href="https://inferenz.ai/contact-us/"><img loading="lazy" decoding="async" class="alignleft wp-image-12051 size-full" style="width: 100%; margin-bottom: 20px;" src="https://inferenz.ai/wp-content/uploads/2025/10/CTA-2-2.gif" alt="" width="1400" height="378" /></a></p>
<h3></h3>
<h3></h3>
<h3><span style="font-weight: 400;">Workflow overview</span></h3>
<div class="table-responsive">
<table>
<tbody>
<tr>
<td><b>Stage</b></td>
<td><b>Process</b></td>
<td><b>Primary Tool</b></td>
<td><b>QA Outcome</b></td>
</tr>
<tr>
<td><b>1</b></td>
<td><span style="font-weight: 400;">Data migration from legacy CRM</span></td>
<td><span style="font-weight: 400;">Migration scripts</span></td>
<td><span style="font-weight: 400;">Source-to-target data movement</span></td>
</tr>
<tr>
<td><b>2</b></td>
<td><span style="font-weight: 400;">Data lake layering</span></td>
<td><span style="font-weight: 400;">Databricks (Bronze → Silver → Gold)</span></td>
<td><span style="font-weight: 400;">Data transformation and enrichment</span></td>
</tr>
<tr>
<td><b>3</b></td>
<td><span style="font-weight: 400;">Automated validation</span></td>
<td><b>Python</b></td>
<td><span style="font-weight: 400;">Record and schema verification</span></td>
</tr>
<tr>
<td><b>4</b></td>
<td><span style="font-weight: 400;">Governance enforcement</span></td>
<td><b>Unity Catalog</b></td>
<td><span style="font-weight: 400;">Role-based access, lineage, and audit logging</span></td>
</tr>
<tr>
<td><b>5</b></td>
<td><span style="font-weight: 400;">Test management</span></td>
<td><b>Zephyr Scale</b></td>
<td><span style="font-weight: 400;">Test execution tracking and reporting</span></td>
</tr>
<tr>
<td><b>6</b></td>
<td><span style="font-weight: 400;">Issue management</span></td>
<td><b>Jira</b></td>
<td><span style="font-weight: 400;">Ticketing, sign-off, and visibility</span></td>
</tr>
</tbody>
</table>
</div>
<p><span style="font-weight: 400;">This structure built confidence through traceability and consistent automation cycles.</span></p>
<h2><span style="font-weight: 400;">Key takeaways from the CRM to D365 transition</span></h2>
<p><img loading="lazy" decoding="async" class="alignleft size-full wp-image-12053" style="width: 100%;" src="https://inferenz.ai/wp-content/uploads/2025/10/Key-Takeaways-from-the-CRM-to-D365-Transition.jpg" alt="" width="1440" height="1029" /></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Treat CRM migration as </span><b>a business transformation</b><span style="font-weight: 400;">, not just data movement.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Use </span><b>Zephyr Scale</b><span style="font-weight: 400;"> for transparent test tracking.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Automate frequent checks using </span><b>Python</b><span style="font-weight: 400;"> to maintain speed and precision.</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Leverage </span><b>Unity Catalog</b><span style="font-weight: 400;"> for governance assurance and compliance.</span></li>
</ul>
<h2><span style="font-weight: 400;">Final thoughts</span></h2>
<p><span style="font-weight: 400;">Migrating to Microsoft Dynamics 365 while building a modern data stack highlighted how deeply </span><b>QA intersects with data engineering and governance</b><span style="font-weight: 400;">.</span></p>
<p><span style="font-weight: 400;">By combining </span><b>Zephyr Scale</b><span style="font-weight: 400;">, </span><b>Python automation</b><span style="font-weight: 400;">, and </span><b>Unity Catalog</b><span style="font-weight: 400;">, we achieved a QA framework that was:</span></p>
<ul>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Structured for traceability,</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Automated for efficiency, and</span></li>
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Governed for compliance.</span></li>
</ul>
<p><span style="font-weight: 400;"><br />
</span><span style="font-weight: 400;">This foundation now serves as a blueprint for future enterprise migrations, ensuring data trust from ingestion to insight.</span></p>
<p>The post <a href="https://inferenz.ai/blogs/qa-in-the-modern-data-stack-using-python-zephyr-scale-unity-catalog-for-end-to-end-quality-assurance/">QA in the Modern Data Stack: Using Python, Zephyr Scale &#038; Unity Catalog for End-to-End Quality Assurance</a> appeared first on <a href="https://inferenz.ai">Inferenz</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
