Uploaded August 2026 | Updated September 2026, 2 weeks ago
DNA data storage has emerged as a promising medium for large-scale, long-term archival preservation. However, existing studies have primarily focused on text, images, and general digital files, while dedicated systems for speech data remain largely unexplored. To fill this gap, we propose VoiceArk, a progressive DNA storage system for robust and scalable speech archiving. As the first stage of the system, we develop a parameterized base layer that employs speech-oriented compression to convert waveform speech into a low-bitrate parameter stream, thereby enabling high compression efficiency, improved error robustness, and intelligible speech reconstruction. To further support progressive access, we envision a deep learning- based enhancement layer that can incrementally improve reconstruction quality as additional sequencing reads become available. Under this design, the system is able to first recover intelligible speech at low read cost and then progressively refine perceptual quality with increased read allocation. Experiments conducted in a simulated DNA channel error environment show that the parameterized base layer achieves intelligible speech reconstruction without error-correcting codes, reaching an STOI score of 0.7 at an error rate of 0.1%. These results demonstrate that the proposed base layer provides a robust speech representation for DNA storage and establishes a practical foundation for progressive speech archiving under a flexible trade-off between read cost and reconstruction quality.
Presented by
Tao Han, Tianjin University
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia
DNA data storage has emerged as a promising medium for large-scale, long-term archival preservation. However, existing studies have primarily focused on text, images, and general digital files, while dedicated systems for speech data remain largely unexplored. To fill this gap, we propose VoiceArk, a progressive DNA storage system for robust and scalable speech archiving. As the first stage of the system, we develop a parameterized base layer that employs speech-oriented compression to convert waveform speech into a low-bitrate parameter stream, thereby enabling high compression efficiency, improved error robustness, and intelligible speech reconstruction. To further support progressive access, we envision a deep learning- based enhancement layer that can incrementally improve reconstruction quality as additional sequencing reads become available. Under this design, the system is able to first recover intelligible speech at low read cost and then progressively refine perceptual quality with increased read allocation. Experiments conducted in a simulated DNA channel error environment show that the parameterized base layer achieves intelligible speech reconstruction without error-correcting codes, reaching an STOI score of 0.7 at an error rate of 0.1%. These results demonstrate that the proposed base layer provides a robust speech representation for DNA storage and establishes a practical foundation for progressive speech archiving under a flexible trade-off between read cost and reconstruction quality.
Presented by
Tao Han, Tianjin University
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia










![DNA MGC+ A Codec for Reliable and Efficient DNA Data Storage
Efficient and reliable data retrieval remains a major challenge in DNA data storage due to the inherent noisiness of the underlying biochemical processes, which lead to both base-level errors and sequence-level dropouts. Here we introduce DNA-MGC+, a novel DNA storage codec designed to enable reliable and efficient data retrieval in the presence of insertion, deletion, and substitution (IDS) errors as well as dropouts. DNA-MGC+ combines an inner coding layer based on the Marker Guess & Check Plus (MGC+) code [1] for correcting IDS errors with an outer Reed-Solomon code that recovers from sequence dropouts and corrects residual inner decoding errors. Our results show that DNA-MGC+ consistently outperforms other codecs across diverse operating conditions. In particular, we observe gains in sequencing depth requirements and decoding time under both Illumina and Nanopore sequencing. We evaluated the performance of DNA-MGC+ in comparison with representative codecs through an in vitro experiment in which sequences encoded using multiple codec configurations were combined in a single oligonucleotide pool. Specifically, a 24-KB compressed file was encoded into oligonucleotides of length 170 bases using two configurations of DNA-MGC+ with different redundancy allocations, design A (1.03 bits/nt) and design B (0.71 bits/nt), as well as two existing codecs: DNA-Aeon [2] (1 bit/nt) and HEDGES [3] (0.61 bits/nt). The oligonucleotide pool was ordered from GenScript (electrochemical synthesis) and sequenced using both Illumina and Oxford Nanopore platforms, with multiple basecalling algorithms evaluated for the Nanopore data. Across all sequencing and basecalling setups, the stored file was recovered with an exact match, albeit with quantitatively different performance outcomes. The results shown in the attached figure indicate that DNA-MGC+ consistently outperforms both DNA-Aeon and HEDGES in terms of the minimum sequencing depth required for reliable decoding, achieving depths below 3x for both Illumina and Nanopore sequencing.
Presented by
Serge Kas Hanna, CNRS
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: https://www.snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: https://snia.org/library
· X: https://twitter.com/SNIA
· LinkedIn: https://linkedin.com/company/snia/ DNA MGC+ A Codec for Reliable and Efficient DNA Data Storage](https://i.ytimg.com/vi/gqRbmqRlTMM/mqdefault.jpg)