The risk of Model Autophagy Disorder: why firms need a Plan B for their data – Opinion
Source: Business Times
Article Date: 22 Jul 2026
Author: Lim Hsin Yin
When AI models are trained repeatedly on content produced by other AI systems, their ability to recognise rare scenarios is usually the first to disappear.
Even as businesses embed artificial intelligence deeper into their daily operations, public content on the Internet is becoming a less reliable source for the context that the technology needs.
Much of the Web’s readily available information has already been heavily mined and fed to large language models (LLMs), and a growing share of what remains is now generated by AI rather than people.
And whenever the quality of the Internet’s information declines, so does the reliability of the AI outcomes built on top of it.
It is in this context that internal corporate documents – records and years of accumulated knowledge – are growing in importance.
Why the Internet is going MAD
Researchers at Rice University call the phenomenon Model Autophagy Disorder, or MAD, while researchers from Oxford and Cambridge have detailed a related problem – model collapse – in a 2024 paper published in Nature.
Both describe what happens when AI models are trained repeatedly on content produced by other AI systems. Their ability to recognise rare scenarios is usually the first to disappear, and these are often exactly the cases where AI intervention is needed most, such as flagging suspicious transactions or regulatory breaches.
This is why organisations are turning to something they already own – their decades of transaction records, customer interactions, compliance documentation and operational data. Much of this has historically been treated as a storage burden, kept to satisfy regulators and rarely touched outside an audit.
That attitude is changing. Feeding AI tools these kind of internal records – data that came from real customers, transactions and employees – rather than leaving them to draw solely on an increasingly AI-saturated public Web, gives those tools a more reliable foundation and grants the business an asset competitors cannot easily replicate.
The same foundation models are available to nearly every organisation, so the real differentiator is the quality and uniqueness of the data behind them.
Yet, few companies are positioned to capitalise on this. Years of piecemeal technology investment have left data scattered across disconnected systems, and many organisations still cannot say with confidence what data they hold or where it resides.
Ensuring corporate data is reliable
Keeping records is not the same as being able to rely on them. Corporate records are only useful as a counterweight to the public Web if they can be proven authentic and unaltered.
That assumption is becoming contested. As internal archives become more central to AI strategy, they also grow in attractiveness to cybercriminals. Instead of locking systems for ransom, attackers increasingly favour quietly tampering with rivals’ records or compromising backups in ways designed to go unnoticed.
This changes what matters for AI readiness. The question is no longer only whether a model is well trained, but whether the data feeding it can still be verified as authentic and how quickly a company can restore a clean version if something goes wrong.
Output can look persuasive and coherent even when the underlying data has been altered, and decisions based on such data will be biased regardless of how sophisticated the underlying LLM is.
Automation raises the stakes further, since processes once subject to layers of human review are now accelerated by AI, compounding errors before anyone catches them.
A regional bank cannot train a fraud-detection model on transaction histories it cannot vouch for. A logistics operator cannot trust inventory data if it cannot be proven untouched.
Regulators are taking note too, expecting organisations to show that AI-driven decisions rest on information that can be traced and verified, not simply retrieved.
Treating records as live infrastructure
As reliable public information grows scarcer, corporate memory may become one of an organisation’s most valuable assets, but only if it is protected. That means treating the archive as live infrastructure, actively monitored and defended, rather than a compliance afterthought.
Getting ready for AI and protecting internal records are not separate agendas. An organisation that cannot vouch for the integrity of its own data has not built a Plan B for the day when models collapse. It has simply moved the same problem in-house.
The writer is vice president, Asean, of Cohesity
Source: The Business Times © SPH Media Limited. Permission required for reproduction.
187