Uploaded August 2026 | Updated September 2026, 2 weeks ago
Motivation: DNA is a promising solution for the global data explosion, yet DNA data storage design remains constrained by the need for sequence orthogonality. Weak, non-specific interactions are often avoided primarily due to lack of models that can accurately and rapidly predict such interactions and lack of large-scale experimental datasets to train such models. This limitation significantly restricts the usable molecular sequence space. Objective: We address these challenges by generating a massive wet lab dataset capturing a broad interaction space and developing a deep learning classifier, BINND, to predict DNA-DNA binding. Method: To build the dataset, a library of randomized 20-mer overhang sequences was annealed to 26 different biotinylated bead primer sequences. Using streptavidin magnetic beads, the bound library strands were separated from the unbound fraction. Post-sequencing and preprocessing, we generated a dataset comprising 144 million 20-mer DNA sequence pairs. BINND is a Convolutional Neural Network classifier trained and evaluated on this large-scale wet lab data. We benchmarked BINND against popular binding predictors such as Gibbs free energy difference, Levenshtein distance, and Hamming distance, focusing on correctness, execution time and memory consumption. We also developed BINND-Lite, a compact version optimized for faster training and lower computational overhead with minimal impact on accuracy. Key results: BINND archives an accuracy of 84%, outperforming the state of the art thermodynamic models by 10% while operating 50x faster. The model generalizes effectively across diverse sequence spaces and varied interaction environments. To demonstrate practical application, we leveraged non-specific interactions to create a physically searchable DNA database representing fictitious children’s book characters and their attributes, enabling fuzzy search capabilities. Conclusion: BINND serves as the first architectural benchmark for DNA-DNA interaction at the 100-million+ sample scale.
Presented by
Gunavaran Brihadiswaran, North Carolina State University
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia
Motivation: DNA is a promising solution for the global data explosion, yet DNA data storage design remains constrained by the need for sequence orthogonality. Weak, non-specific interactions are often avoided primarily due to lack of models that can accurately and rapidly predict such interactions and lack of large-scale experimental datasets to train such models. This limitation significantly restricts the usable molecular sequence space. Objective: We address these challenges by generating a massive wet lab dataset capturing a broad interaction space and developing a deep learning classifier, BINND, to predict DNA-DNA binding. Method: To build the dataset, a library of randomized 20-mer overhang sequences was annealed to 26 different biotinylated bead primer sequences. Using streptavidin magnetic beads, the bound library strands were separated from the unbound fraction. Post-sequencing and preprocessing, we generated a dataset comprising 144 million 20-mer DNA sequence pairs. BINND is a Convolutional Neural Network classifier trained and evaluated on this large-scale wet lab data. We benchmarked BINND against popular binding predictors such as Gibbs free energy difference, Levenshtein distance, and Hamming distance, focusing on correctness, execution time and memory consumption. We also developed BINND-Lite, a compact version optimized for faster training and lower computational overhead with minimal impact on accuracy. Key results: BINND archives an accuracy of 84%, outperforming the state of the art thermodynamic models by 10% while operating 50x faster. The model generalizes effectively across diverse sequence spaces and varied interaction environments. To demonstrate practical application, we leveraged non-specific interactions to create a physically searchable DNA database representing fictitious children’s book characters and their attributes, enabling fuzzy search capabilities. Conclusion: BINND serves as the first architectural benchmark for DNA-DNA interaction at the 100-million+ sample scale.
Presented by
Gunavaran Brihadiswaran, North Carolina State University
This is a presentation from the 2026 Storage and Computing with DNA Conference.
· Learn More about the SNIA DNA Data Storage Alliance: snia.org/groups/snia-dna-technology-affiliate
· SNIA Educational Library: snia.org/library
· X: twitter.com/SNIA
· LinkedIn: linkedin.com/company/snia




Learn more about SAS and the STA community: https://www.snia.org/sta
#shorts
Learn More about SNIA and standards:
• SNIA Website: https://snia.org/
• SNIA Educational Library: https://snia.org/library
• X: https://x.com/SNIA
• LinkedIn: https://www.linkedin.com/company/snia/ Short Video: SAS Continues to Innovate with SBC-5](https://i.ytimg.com/vi/ae_AT_JHIqc/mqdefault.jpg)






