Unlock a world of possibilities! Login now and discover the exclusive benefits awaiting you.
There’s a moment every data team knows. An executive questions a number in a dashboard. An AI model returns a recommendation that doesn’t make sense. A downstream report shows a discrepancy no one can explain. The root cause turns out to be a data quality issue that has been sitting in the pipeline for weeks.
The frustrating part is that the issue rarely originated where it was found. It perhaps started upstream, from an incorrect transformation, a schema change that wasn’t handled, or a source system that began emitting invalid values two sprints ago. Every unidentified data issue compounds downstream, driving rework, slowing decisions, and eroding trust in analytics and AI. Gartner research estimates that the cost of poor data quality averages $12.9 million per organizatio...
By the time an AI app or a business user surfaces the problem, it has aged, compounded, and become harder to trace. There is a better containment model, and software engineering figured it out years ago.
In software development, “shifting left” means moving quality and governance checks earlier in the lifecycle. Barry Boehm’s research quantified why it matters for code development. A defect caught during development costs the least to fix, a defect found during testing costs considerably more, and a defect that reaches production costs the most by far. The multiplier compounds at every stage you move downstream, because context degrades, blast radius widens, and the fix becomes harder to isolate. Software engineering recognized the problem years ago, internalized it and reorganized the whole discipline around it. This gave rise to DevOps Research and Assessment (DORA) framework for continuously improving software delivery.
Furthermore, the DORA concept can be applied to other domains beyond just software engineering. Indeed, we can apply it to data engineering as well, specifically data quality. Thinking through the lens of DORA metrics, KPIs such as deployment frequency, lead time for changes, change failure rate, and mean time to recovery can be mapped to equivalents within data quality engineering.
|
DORA Metric |
Software Engineering |
Data Quality Equivalent |
|
Deployment Frequency |
How often code is deployed |
How often quality rules and controls are deployed into live pipelines |
|
Lead Time for Changes |
Time from commit to production |
Time from identifying a DQ requirement to enforcing the rule in production |
|
Change Failure Rate |
% of deployments causing incidents |
% of pipeline or schema changes that introduce data defects |
|
Mean Time to Recovery |
Time to restore service |
Time to detect, remediate, and validate a data quality issue |
Just like code, data follows the same curve. A quality issue caught at ingestion costs a fraction of what it costs when it surfaces in a dashboard, a regulatory report, or an AI model.
Shifting quality left has a specific meaning for data engineers. Validation logic runs at the point where data is created and transformed, not the point where it is consumed. When a new source is onboarded, completeness and schema checks run immediately. When a transformation is written, it carries assertions about what the output should look like. When a pipeline runs, anomaly detection fires in real time, not three weeks later when someone notices a metric looks off. This repositions the data engineer from pipeline builder to quality-accountable producer.
The data engineering role itself is changing in the process. Hand-crafting rule expressions and wiring checks into every job used to fill the day. Qlik’s Agentic Data Engineering capabilities, now generally available in Qlik Talend Cloud changes what that work looks like. With the Data Quality Agent, an engineer describes what good data looks like in plain language; the agent finds the target dataset, checks existing fields and rules to avoid redundancy, generates the rule expressions, and lets the engineer refine them before anything is written. The engineer directs and reviews; the agent does the mechanical work. Building pipelines and fixing data quality issues becomes a higher-leverage discipline allowing the quality coverage to widen well beyond the specialists who used to own it.
Shifting quality left only works if you can see what is happening across your pipelines. This is where data observability comes in, and it is worth being precise because it is often conflated with monitoring.
Monitoring tells you that something failed. Observability tells you why, when it started, and where it originated. A monitoring check might flag that a table has unexpected nulls. An observability signal tells you that null values in a specific column rose 40% over the last two hours, that the spike correlates with a transformation updated yesterday, and that this pattern has appeared twice before in the past month. That is a signal you can act on.
Effective observability means instrumenting pipelines continuously for freshness (is data arriving when it should?), completeness (are expected volumes showing up?), schema consistency (are structure changes propagating correctly?), distribution drift (are values behaving as they historically have?), and rule conformance (are business-defined constraints holding?). These run as ongoing signals that are continuously woven into the operational fabric.
The Qlik Trust Score for AI condenses the trustworthiness of a dataset or data product into a single, intuitive measure. Because Trust Score is tracked over time, an engineer can compare it before and after the Data Quality Agent acts — seeing exactly which signals moved the score and by how much. The evolution of Trust Score becomes the feedback loop: every new rule, every remediation, and every schema change shows up as a measurable shift, so the impact of each intervention is visible rather than inferred. When you combine that visibility with earlier enforcement, you get a compounding effect. Rules embedded in pipelines generate signals upstream. These signals surface issues while they are still close to the source. Resolution happens faster because the context is fresh. Over time, you build a map of where your data is fragile and where it is reliable, and you address quality proactively rather than reactively.
The technical model only works if the organizational model supports it. Traditional data quality programs concentrate ownership in a small governance team operating at a remove from the engineering work. That team defines policy. Engineers implement pipelines. Governance audits outputs. The feedback loop is long, and by the time an issue surfaces through that channel, it has already caused downstream harm.
The old stereotype was governance as the data police: no, you can’t use that field, please attend next month’s council. Even well-run enablement teams operated from the outside, persuading and reminding everyone else to comply. That does not scale. Quality and governance that survives AI shows up where the work happens, in pull requests, in CI/CD, and in the pipelines themselves.
A left-shifted model distributes accountability differently. Data engineers own quality at the point of production. Domain teams are responsible for the data they generate. Business stewards contribute the contextual judgment automated systems cannot provide, because they know what “correct” means for their domain and can tell a real problem from expected variance.
This is a change in how quality work is structured and who participates, not a reorganization. Automated checks handle the mechanical layer. Human judgment handles the interpretation layer. Two more of the new agents in Qlik Talend Cloud make that split workable. Data products are how the gap between data producers and data consumers finally closes — producers publish a governed, contract-bound asset; consumers get something discoverable, documented, and safe to trust. The Data Product Agent manages the full lifecycle of a data product through conversation, with pre-activation validation, dependency awareness, and explicit confirmation before anything is created, deactivated, or deleted. Producers evolve the product without breaking downstream consumers, and governance is enforced at every step without adding process overhead, yet allowing human judgment to be in the loop at each decision point. The Catalog and Business Glossary Agent then makes those products understandable in business terms. It derives glossary terms from business domain context, data product structure, field metadata, and app measures and dimensions, and links each term to the assets it describes. Documentation stays in sync with the data instead of rotting the moment someone renames a column, and other agents inherit correct, governed metadata to reason over.
The Data Quality, Catalog and Business Glossary, and Data Product Agents, along with Declarative Pipelines, are all generally available today in Qlik Talend Cloud. Clive Bearman walks through each one, with demos, in the full launch announcement: From Intent to Trusted Data: Agentic Data Engineering is Now Generally Available. If you are ready to move quality upstream, start here:
🔗 Release notes: here
🔗 Want a demo, join the session : here
You must be a registered user to add a comment. If you've already registered, sign in. Otherwise, register and sign in.