Data governance frameworks, including RBI's new draft, define data risk as arising from deficiencies in data or its management. Cybersecurity frameworks define cyber risk as the risk of unauthorised access, disclosure, disruption, or destruction of information assets. These definitions evolved within different disciplines, owned by different teams and governed through different board committees. Historically that separation was workable because data quality failures and cyber incidents were usually distinct problems with distinct remediation playbooks. It no longer is, and the reason is structural.
Integrity controls were built for error, not attack
RBI's draft operationalises integrity as protection against unauthorised modification, corruption, or loss, much as BCBS 239 and most enterprise data governance models do. When these frameworks were first written, threats to integrity were almost entirely accidental: a corrupted file, a botched migration, an operator error, a batch job that dropped records. The controls built for that era, validation rules, reconciliation, change logs, and access entitlements tied to job function, are well suited to catching accidents, because accidents tend to be inconsistent. An error in one place rarely covers its tracks everywhere else.
What has changed is not the definition of integrity but the threat model that definition now has to encompass. The distinction is not between internal and external failures. It is between failures that reveal themselves and failures deliberately engineered to remain invisible. Traditional data governance evolved to detect the former. Modern data integrity depends as much on resisting deliberate manipulation as on detecting accidental error.
A stolen credential can make malicious modifications appear authorised. A compromised upstream supplier can inject records that are internally well formed and reconcile perfectly with downstream controls. An attacker who understands the validation rules can manipulate data specifically to satisfy them. In each case the data satisfies every internal governance check while its trustworthiness has already been lost.
None of this means data governance frameworks ignore cybersecurity. RBI's draft, like BCBS 239, explicitly names information security and data privacy alongside data quality and lifecycle management. The gap is not recognition, it is ownership: accountability for data integrity ends before accountability for the controls capable of preserving that integrity begins.
Accountability for data integrity ends before accountability for the controls capable of preserving that integrity begins.
What a breach actually costs, beyond disclosure
The conventional framing of a breach is a confidentiality failure: data that should have stayed private did not. That framing is complete for the cases the industry has spent two decades building controls against: card data, customer records, credentials. It is not complete when the point of the intrusion is not to exfiltrate data but to alter it, degrade it, or seed it with content that will be trusted and acted on downstream, poisoning the data lake rather than emptying it. The harm is that something the regulated entity (RE) relies on to make decisions has quietly stopped being true, and a governance framework built to catch internal deficiencies has no reason to look for it.
This is why cyber resilience cannot be handled by cross-reference. A data governance framework that says, in effect, see the separate cybersecurity policy for anything involving an external actor, is drawing the boundary in exactly the place a sophisticated intrusion is designed to exploit. Governance theory treats accountability without corresponding authority as a structural flaw, because obligations cannot be discharged where the controls determining success belong to another function. That is precisely the position today of the Data Owner, the business executive accountable under the governance framework for the quality, integrity, availability, and appropriate use of a defined data domain: accountable for the integrity of data in their domain, without authority over the security controls that determine whether that integrity survives external compromise.
Where AI makes the gap expensive rather than theoretical
AI makes the consequences materialise faster and become harder to reverse.
Retrieval-augmented systems make this immediate rather than hypothetical. These systems answer questions by pulling from a knowledge repository, a vector store built from the RE's own documents and data, at the moment of the query, without the underlying model ever being retrained. Poison that repository with a small number of records engineered to be retrieved for specific queries, and the system begins producing incorrect or manipulated outputs the next time someone asks a relevant question. No training run is disturbed and no model weight changes. The only thing that changed is the population feeding retrieval, and that population looks exactly like the kind of internally well formed, reconciled, complete dataset a framework built to catch accidents has no reason to flag.
Fine-tuning is the harder version of the same problem, and while it is not yet common practice at most REs, it is the direction most are already moving in. A poisoned retrieval store can be re-indexed once the poisoning is found. A model fine-tuned on the same compromised data cannot be corrected the same way: the alteration is no longer sitting in a record that can be removed or a row that can be restored, it is distributed across model parameters with no practical way of isolating or removing the poisoned influence. Remediation ultimately requires retraining, which assumes the RE even knows which training run needs to be redone.
What this means in practice
The fix is not a new committee or a new framework bolted alongside the existing two. It requires redrawing the boundary of the Data Owner's accountability, and underneath that, a redefinition of data risk itself. The current definition, risk arising from deficiencies in data or in its management, is an internal category, covering errors, gaps, and mismanagement the RE itself is responsible for. It needs to widen to include risk arising from external adversarial compromise affecting the trustworthiness, availability, or integrity of data: action taken deliberately by someone outside the organisation's control, not merely a mistake made within it. Once data risk includes both, the Data Owner's mandate must include both, and the mapping of roles and responsibilities must give the Data Owner real line of sight into the cyber controls protecting their data, not a document they do not own.
The logical consequence is that cyber resilience cannot remain merely a referenced dependency of data governance. It must become one of the governance controls through which Data Owners discharge their accountability for integrity. A governance framework that separates responsibility for data integrity from responsibility for defending that integrity against deliberate compromise creates a structural gap in accountability that no amount of coordination between functions can close. For data governance frameworks, including RBI's draft, the implication is straightforward: cyber resilience is no longer simply a supporting control domain. It has become part of the governance architecture required to assure the integrity of data itself.