Uploaded July 2026 | Updated September 2026, 6 hours ago
What does the AI storage stack actually look like — and why does it matter?
At Computex 2026, WD Chief Product Officer Ahmed Shihab breaks down the architecture powering today's AI infrastructure. Drawing on 20 years of experience at NetApp, AWS, Microsoft, and WD, Ahmed explains why AI creates a fundamentally different storage challenge than anything the industry has seen before.
In this session, he covers:
Why AI is two compute workloads (training and inference) but three storage workloads
How inference output — not just training data — is one of the fastest-growing drivers of storage demand
A live example showing how generating an 8-second video required reading 40GB of data and produced 320GB of intermediate data across seven attempts
Why the AI storage stack has more layers than the cloud, not fewer — from HBM and GPU memory down to flash caches, vector/embedding stores, and bulk HDD storage
The economics of scale: at 200 exabytes, choosing all-flash over a balanced HDD/flash architecture costs the equivalent of 3 million GPUs
IDC's forecast of 300 zettabytes of additional AI-driven storage demand by the end of the decade
Ahmed makes the case that the 80/20 balance of HDD to flash in AI infrastructure is elegant, pragmatic, and economically necessary — and that data, not compute, is the foundation everything else is built on.
What does the AI storage stack actually look like — and why does it matter?
At Computex 2026, WD Chief Product Officer Ahmed Shihab breaks down the architecture powering today's AI infrastructure. Drawing on 20 years of experience at NetApp, AWS, Microsoft, and WD, Ahmed explains why AI creates a fundamentally different storage challenge than anything the industry has seen before.
In this session, he covers:
Why AI is two compute workloads (training and inference) but three storage workloads
How inference output — not just training data — is one of the fastest-growing drivers of storage demand
A live example showing how generating an 8-second video required reading 40GB of data and produced 320GB of intermediate data across seven attempts
Why the AI storage stack has more layers than the cloud, not fewer — from HBM and GPU memory down to flash caches, vector/embedding stores, and bulk HDD storage
The economics of scale: at 200 exabytes, choosing all-flash over a balanced HDD/flash architecture costs the equivalent of 3 million GPUs
IDC's forecast of 300 zettabytes of additional AI-driven storage demand by the end of the decade
Ahmed makes the case that the 80/20 balance of HDD to flash in AI infrastructure is elegant, pragmatic, and economically necessary — and that data, not compute, is the foundation everything else is built on.










