Uploaded May 2026 | Updated September 2026, 2 weeks ago
The rapid acceleration of AI model complexity, scale, and diversity has exposed a widening gap between real world AI storage workloads and the benchmarks traditionally used to evaluate storage systems. As AI evolves from large scale training to retrieval augmented generation, vector databases, and KV cache intensive inference — storage systems face new I/O patterns, new bottlenecks, and new expectations for determinism, bandwidth, and latency. Accurately representing these behaviors in a benchmark is now as challenging as architecting the storage itself.
This session explores the emerging landscape of AI focused storage benchmarking, with emphasis on two major industry efforts: MLPerf Storage v3.0, the next iteration of MLCommons’ benchmark suite, and the newly forming SNIA AI Data Workloads Technical Working Group (TWG). Together, these groups aim to close the gap between how AI workloads behave in production and how storage devices are evaluated in labs.
We will outline the core challenges in representing AI behavior at the storage layer and identify the specific pain points storage vendors and system architects face when attempting to reproduce realistic AI I/O, and why traditional metrics such as throughput and latency from 4-corners synthetic tests fail to reflect true system performance.
The session will also introduce who in the industry is working to solve this problem and how:
• MLPerf Storage v3.0 efforts to incorporate new AI models, synthetic dataset generators, and expanded training/inference pipelines.
• SNIA’s AI Data Workloads TWG, focused on defining open, standardized approaches for characterizing AI storage behavior across retrieval, training, inference, vector DBs, KV cache management, and GPU initiated I/O.
As AI workloads continue to evolve at unprecedented speed, so too must the benchmarks that evaluate storage solutions. This talk provides the roadmap for how the industry is rising to that challenge, and how attendees can participate in shaping the next era of AI aligned storage benchmarking.
Presented by Wes Vaske | Micron Technology - Senior Member of Technical Staff
Learn More:
• SDC: StorageAI Website: snia.org/sniadeveloper/storageai
• SNIA Website: snia.org
• SNIA Educational Library: snia.org/library
• X: twitter.com/SNIA
• LinkedIn: linkedin.com/company/snia
The rapid acceleration of AI model complexity, scale, and diversity has exposed a widening gap between real world AI storage workloads and the benchmarks traditionally used to evaluate storage systems. As AI evolves from large scale training to retrieval augmented generation, vector databases, and KV cache intensive inference — storage systems face new I/O patterns, new bottlenecks, and new expectations for determinism, bandwidth, and latency. Accurately representing these behaviors in a benchmark is now as challenging as architecting the storage itself.
This session explores the emerging landscape of AI focused storage benchmarking, with emphasis on two major industry efforts: MLPerf Storage v3.0, the next iteration of MLCommons’ benchmark suite, and the newly forming SNIA AI Data Workloads Technical Working Group (TWG). Together, these groups aim to close the gap between how AI workloads behave in production and how storage devices are evaluated in labs.
We will outline the core challenges in representing AI behavior at the storage layer and identify the specific pain points storage vendors and system architects face when attempting to reproduce realistic AI I/O, and why traditional metrics such as throughput and latency from 4-corners synthetic tests fail to reflect true system performance.
The session will also introduce who in the industry is working to solve this problem and how:
• MLPerf Storage v3.0 efforts to incorporate new AI models, synthetic dataset generators, and expanded training/inference pipelines.
• SNIA’s AI Data Workloads TWG, focused on defining open, standardized approaches for characterizing AI storage behavior across retrieval, training, inference, vector DBs, KV cache management, and GPU initiated I/O.
As AI workloads continue to evolve at unprecedented speed, so too must the benchmarks that evaluate storage solutions. This talk provides the roadmap for how the industry is rising to that challenge, and how attendees can participate in shaping the next era of AI aligned storage benchmarking.
Presented by Wes Vaske | Micron Technology - Senior Member of Technical Staff
Learn More:
• SDC: StorageAI Website: snia.org/sniadeveloper/storageai
• SNIA Website: snia.org
• SNIA Educational Library: snia.org/library
• X: twitter.com/SNIA
• LinkedIn: linkedin.com/company/snia








![Nanopore sequencing of synthetic libraries of RNA oligonucleotides
Photolithography is one of the very approaches that allow for the synthesis of nucleic acid microarrays in situ, and characteristic aspects of in situ microarray synthesis are high-throughput and high-density, delivering several hundreds of thousands of unique sequences in a single run and on a single, small surface (Figure 1). Microarray synthesis has traditionally focused on the preparation of DNA microarrays to obtain complex DNA libraries. These have been used in the context of DNA data storage, gene synthesis and other nanotechnology applications [1]. Recently, our group has shown that photolithography is amenable to prepare RNA microarrays as well, at identical throughput and density [2]. It remains the only available chemical approach that can deliver complex synthetic RNA libraries with total control on the sequence. RNA microarrays can be used to interrogate the sequence preference of enzymes and RNA-binding proteins, but they are also ideally poised to generate RNA libraries for off-array applications. We can produce pools of RNA sequences between 75 and 100-nt in length which can be sequenced directly by Nanopore sequencing without any intermediate purification step [3]. Our photolithography platform also allows for the introduction of biologically relevant base modifications, of which m6A, 5mC and inosine are already available and preliminary data shows that m6A can be accurately basecalled. Simultaneously, nanopore sequencing data returns crucial information on the synthetic error-rate of RNA photolithography. This talk will focus on presenting the technology of RNA photolithography and on describing how RNA libraries can be prepared and sequenced.
Presented by
Jory Lietard, University of Vienna
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: https://www.snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: https://snia.org/library
· X: https://twitter.com/SNIA
· LinkedIn: https://linkedin.com/company/snia/ Nanopore sequencing of synthetic libraries of RNA oligonucleotides](https://i.ytimg.com/vi/VNJYQbz7MTY/mqdefault.jpg)

