Post 5: Data Quality Is Not Data Cleansing… Stop Confusing Them

THE ENTERPRISE DATA GOVERNANCE PLAYBOOK — Post 5 of 13

Data Quality Is Not Data Cleansing — Stop Confusing Them

By Greg Briscoe, Senior Solution Architect — Enterprise Data Management

 

If I could eliminate one misconception from enterprise data management, it would be this. Data quality and data cleansing are the same thing. They are not. And confusing them is costing your organization millions in reactive remediation that never addresses root causes.

I’ve had this conversation in boardrooms, in project war rooms, and in vendor evaluation sessions. “We need data quality” almost always translates to “we need to clean up the data we already have.” That’s data cleansing. It’s a necessary activity. But it is not a quality strategy. Treating cleansing as quality is like treating the emergency room as your primary care physician you’ll survive each crisis, but you’ll never get healthy.

 

The Cleansing Trap

Here’s how most organizations handle data quality today. Data arrives from source systems in varying states of completeness, accuracy, and consistency. It flows through an ETL layer where scrubbing rules remove duplicates, standardize formats, and patch known issues. The cleaned data lands in the warehouse or reporting environment. Reports look reasonable. Everyone exhales.

Then next quarter, the same data arrives from the same source systems with the same quality issues. The same ETL rules fire. The same scrubbing happens. The same cycle repeats. Quarter after quarter. Year after year. Forever.

This is the cleansing trap, and it’s insidious because it feels productive. Things are happening. Data is being fixed. Reports are being generated. But nothing is actually improving. The root causes the reasons the data was wrong in the first place are never addressed. The source systems keep producing the same defective data because nobody has changed the process, enforced the rules, or assigned accountability at the point of origin.

Meanwhile, the ETL team grows. The scrubbing rules multiply. The edge cases accumulate. And the organization pays an escalating annual cost to remediate symptoms while the disease progresses untreated.

 

Quality at the Source

The alternative philosophy is straightforward but requires a fundamental shift in organizational thinking: fix data quality issues at the source, not downstream.

Data quality is a much more involved process that focuses an organization’s resources on addressing quality issues at the source versus after the fact. It is far easier and cheaper to fix data issues at the source than to rely completely on data cleansing and scrubbing. Ownership and accountability for data quality must reside with the business owners of source systems.

This is a governance statement, not a technology statement. You can’t buy a tool that fixes quality at the source. You need a process that identifies quality failures, traces them back to their origin, assigns accountability to the person or system that introduced the defect, and changes the process so that defect can’t recur. That’s quality management. Everything else is just cleaning up.

 

The Fundamental Difference Data cleansing asks: “How do we fix this data?” Data quality asks: “Why was this data wrong and how do we ensure it’s right the first time?”

What Quality Actually Looks Like

Data quality is not a separate, single activity bolted onto the side of your data management process. Data quality is incorporated into every aspect of the governance framework. The prime function of governance is to improve and maintain the quality of the data thus, to be successful at governance, quality must be continuously measured and the results continuously fed back into the governance process.

In practice, this means:

  • Measurement: Every data domain has defined quality metrics completeness, accuracy, consistency, timeliness that are measured automatically and reported continuously
  • Feedback: Quality measurements flow back to data owners and stewards as actionable insights, not just dashboards that nobody reads
  • Root cause resolution: When quality drops, the response targets the source process, not the downstream remediation
  • Prevention: Business rules enforced at entry prevent quality defects from entering the system in the first place
  • Accountability: Every data element has a named owner responsible for its quality, and quality metrics are part of that owner’s operational scorecard

 

This is a fundamentally different operating model than “clean it in the ETL.” It requires organizational commitment, executive sponsorship, and a governance framework that treats quality as a continuous operational discipline not a quarterly cleanup project.

 

How EDM Enforces This

Oracle EDM Cloud operationalizes this philosophy through several mechanisms that make quality at the source the default rather than the exception:

  • Business rule enforcement at entry: Most of the data record will be derived, defaulted, or inherited preventing manual data-entry errors. Validation rules fire when data is entered or changed, preventing defective records from entering the governed repository.
  • Request-driven recorded actions: Every change flows through a governed workflow that captures who requested it, who approved it, and what validation rules were applied
  • Profiling and monitoring: Continuous data quality measurement against defined metrics, with alerts when quality degrades
  • Validation against source systems: Cross-system reconciliation ensures that the governed repository remains aligned with all consuming applications
  • Role-based access with quality gates: Only authorized users can make changes, and only after those changes pass all applicable quality validations

 

The technology makes quality at the source operationally feasible, but the technology without the governance philosophy is just an expensive gate. You need both the organizational commitment to prevent defects at origin, and the platform that enforces that commitment systematically.

 

Stop budgeting for annual data cleanup projects. Start investing in governance that prevents the mess from happening. The best data quality strategy is the one where the cleansing budget goes to zero not because you stopped caring, but because you stopped needing it.

 

Next: Why trying to govern all your master data the same way is the second most expensive mistake.

 

Greg Briscoe is a Senior Solution Architect specializing in Oracle EPM, EDM, DRM, ERP, master data governance, and large-scale transformation programs. With experience spanning hundreds of enterprise engagements, he helps organizations design and operationalize data governance capabilities that outlast individual projects and compound in value with every transformation initiative.

 

Add Comment