Uploaded August 2026 | Updated September 2026, 2 weeks ago
We discuss our recent progress using next-generation sequencing (NGS) analysis methods to characterize oligo pools created by our complementary metal-oxide- semiconductor (CMOS) chip-based ultra-high-throughput DNA synthesis platform. Typical synthetic biology pools consist of tens of thousands to a million unique sequences (104 - 106 ). Using phosphoramidite chemistry with electrochemical deblocking, we have demonstrated successful parallelized synthesis of 15 x 106 unique 84mer sequences in a single run - the largest per-run throughput capability of which we are aware. The internal data payload of this pool had an estimated total error rate of 0.5% per base and a per-sequence yield uniformity such that the 90% spread is within a 5-fold range (P95/P05 = 5). This oligo quality allowed for lossless recovery of a 15- megabyte archive encoded in the pool, translating to an estimated write rate of 2 kilobytes per second. Processing such high-complexity reference libraries essential for DNA data storage poses unique computational and analytical challenges, not only during the decode step of an archival workflow, but also during the oligo quality assessment stage of development. In this talk, we present methods for assessing error introduction and shifts in copy number representation using standard bioinformatics toolkits in an NGS analysis pipeline, in particular using the short-read sequence aligners Bowtie2 and minimap2. As the size of the oligo pool increases both in length and number of sequences, these aligners must be tuned to balance processing speed and mapping sensitivity while maintaining accurate identification of the error profiles.
Presented by
Rodney Agayan, Atlas Data Storage
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia
We discuss our recent progress using next-generation sequencing (NGS) analysis methods to characterize oligo pools created by our complementary metal-oxide- semiconductor (CMOS) chip-based ultra-high-throughput DNA synthesis platform. Typical synthetic biology pools consist of tens of thousands to a million unique sequences (104 - 106 ). Using phosphoramidite chemistry with electrochemical deblocking, we have demonstrated successful parallelized synthesis of 15 x 106 unique 84mer sequences in a single run - the largest per-run throughput capability of which we are aware. The internal data payload of this pool had an estimated total error rate of 0.5% per base and a per-sequence yield uniformity such that the 90% spread is within a 5-fold range (P95/P05 = 5). This oligo quality allowed for lossless recovery of a 15- megabyte archive encoded in the pool, translating to an estimated write rate of 2 kilobytes per second. Processing such high-complexity reference libraries essential for DNA data storage poses unique computational and analytical challenges, not only during the decode step of an archival workflow, but also during the oligo quality assessment stage of development. In this talk, we present methods for assessing error introduction and shifts in copy number representation using standard bioinformatics toolkits in an NGS analysis pipeline, in particular using the short-read sequence aligners Bowtie2 and minimap2. As the size of the oligo pool increases both in length and number of sequences, these aligners must be tuned to balance processing speed and mapping sensitivity while maintaining accurate identification of the error profiles.
Presented by
Rodney Agayan, Atlas Data Storage
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia





