Uploaded August 2026 | Updated September 2026, 2 weeks ago
Assembly-based DNA synthesis is emerging as a promising strategy to scale DNA data storage. However, this approach introduces structural constraints—such as fragment overlaps, motif restrictions, and length limits—that fundamentally reshape the encoding problem. These constraints influence not only how digital information is translated into DNA sequences, but also how fragments must be organized to support efficient molecular assembly and reliable data recovery.
In this work, we extend the proprietary end-to-end simulator DNAssim to support assembly-based synthesis workflows. The framework explicitly models assembly constraints during simulation and integrates assembly-aware encoding and validation procedures. In addition, we introduce novel clustering algorithms designed to organize DNA fragments while respecting assembly requirements. Clustering plays a central role in this process: fragments must be grouped in ways that preserve valid overlaps, avoid problematic sequence motifs, and maintain compatibility with downstream assembly protocols.
By capturing the interaction between encoding constraints, fragment organization, and assembly requirements, DNAssim enables systematic exploration of assembly-aware DNA storage designs. The framework allows researchers to evaluate how fragment structure, overlap strategies, and clustering approaches influence overall system performance. In particular, it provides a flexible platform to study how fragment organization impacts scalability, storage efficiency, and robustness during both synthesis and decoding.
Presented by
Alessia Marelli, Avaneidi
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia
Assembly-based DNA synthesis is emerging as a promising strategy to scale DNA data storage. However, this approach introduces structural constraints—such as fragment overlaps, motif restrictions, and length limits—that fundamentally reshape the encoding problem. These constraints influence not only how digital information is translated into DNA sequences, but also how fragments must be organized to support efficient molecular assembly and reliable data recovery.
In this work, we extend the proprietary end-to-end simulator DNAssim to support assembly-based synthesis workflows. The framework explicitly models assembly constraints during simulation and integrates assembly-aware encoding and validation procedures. In addition, we introduce novel clustering algorithms designed to organize DNA fragments while respecting assembly requirements. Clustering plays a central role in this process: fragments must be grouped in ways that preserve valid overlaps, avoid problematic sequence motifs, and maintain compatibility with downstream assembly protocols.
By capturing the interaction between encoding constraints, fragment organization, and assembly requirements, DNAssim enables systematic exploration of assembly-aware DNA storage designs. The framework allows researchers to evaluate how fragment structure, overlap strategies, and clustering approaches influence overall system performance. In particular, it provides a flexible platform to study how fragment organization impacts scalability, storage efficiency, and robustness during both synthesis and decoding.
Presented by
Alessia Marelli, Avaneidi
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia


![Nanopore sequencing of synthetic libraries of RNA oligonucleotides
Photolithography is one of the very approaches that allow for the synthesis of nucleic acid microarrays in situ, and characteristic aspects of in situ microarray synthesis are high-throughput and high-density, delivering several hundreds of thousands of unique sequences in a single run and on a single, small surface (Figure 1). Microarray synthesis has traditionally focused on the preparation of DNA microarrays to obtain complex DNA libraries. These have been used in the context of DNA data storage, gene synthesis and other nanotechnology applications [1]. Recently, our group has shown that photolithography is amenable to prepare RNA microarrays as well, at identical throughput and density [2]. It remains the only available chemical approach that can deliver complex synthetic RNA libraries with total control on the sequence. RNA microarrays can be used to interrogate the sequence preference of enzymes and RNA-binding proteins, but they are also ideally poised to generate RNA libraries for off-array applications. We can produce pools of RNA sequences between 75 and 100-nt in length which can be sequenced directly by Nanopore sequencing without any intermediate purification step [3]. Our photolithography platform also allows for the introduction of biologically relevant base modifications, of which m6A, 5mC and inosine are already available and preliminary data shows that m6A can be accurately basecalled. Simultaneously, nanopore sequencing data returns crucial information on the synthetic error-rate of RNA photolithography. This talk will focus on presenting the technology of RNA photolithography and on describing how RNA libraries can be prepared and sequenced.
Presented by
Jory Lietard, University of Vienna
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: https://www.snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: https://snia.org/library
· X: https://twitter.com/SNIA
· LinkedIn: https://linkedin.com/company/snia/ Nanopore sequencing of synthetic libraries of RNA oligonucleotides](https://i.ytimg.com/vi/VNJYQbz7MTY/mqdefault.jpg)







