Home
Storage

Synthetic Data: Storage's Next Growth Driver

AI models don't just consume data. They manufacture it.

Gartner chart projecting synthetic data will overshadow real data in AI models by 2030
Pedro CoutoJul 14, 20262 min read
Share
16 reads

Context changes constantly. Every shift in context forces a new generation of tokens. Each token can represent a new interpretation, a new line of reasoning, a new decision.

We are entering a new industrial era. Large enterprises will produce intelligence at scale, around the clock, every day of the year.

Most of what these processes generate gets stored. This is synthetic data. It is not a byproduct. It is the primary output of automated reasoning at scale.

Some practitioners use a narrower definition of synthetic data, limiting it to content created specifically to train other models, such as generated medical images, simulated financial transactions, artificial customer conversations, synthetic system logs, and virtual driving scenarios.

A broader view includes anything a model authors for a human to read: a report, a summary, a transcript. Storage does not care about the distinction. Both consume capacity.

Why storage teams should care

Gartner projects that by 2030, synthetic data will overshadow real data in AI model training. Real data grows linearly. Synthetic data grows exponentially, generated by models that never stop running.

Every inference cycle writes to storage. Every retry writes to storage. Every reasoning trace writes to storage. None of this is optional storage. It is the exhaust of intelligence production, and it has to land somewhere.

What this means for capacity planning

Traditional growth models assume human-generated data: documents, logs, transactions. Synthetic data breaks that assumption. Volume no longer tracks headcount or user activity. It tracks compute.

Plan for it now. The organizations that treat synthetic data as a storage tier problem today will not be scrambling for capacity in 2030.

About the author

NetApp A-Team member focused on enterprise storage and AI infrastructure

Pedro Couto · Enterprise Infrastructure Architect

Pedro works at the intersection of NetApp ONTAP, hybrid cloud, data protection and high-performance AI platforms. Designing and delivering enterprise infrastructure for organizations where downtime is not an option. A recognized member of the NetApp A-Team and holder of the NetApp Subject Matter Expert Elite designation, he is one of the authors of many certifications of NetApp such as NCDA, NCIE-DP, NCIE-MetroCluster, Storage Engineer and more.

Comments

0

No comments yet. Start the conversation.

Up to 2,000 characters. Plain text only.