Uploaded April 2016 | Updated September 2026, 2 days ago
This talk was given by undergraduate Jake Bogerd during the 10th Annual Computer Science Undergraduate Research Symposium in 2016. Jake‘s research was supervised by Dr. Jan Prins.
“A Method for Construction of a Splice Graph from RNA Sequence Data”
Although genetic information is stored in DNA, the RNA transcripts obtained from DNA indicate which genes are actually active in any given cell. Modern high-throughput sequencing techniques allow accurate sequencing of short RNA fragments, which may then be aligned to a reference genome. These alignments can be summarized by constructing a “splice graph”, in which nodes represent genomic coordinates and edges represent sequences that are retained or spliced out of observed transcripts. Each full-length transcript corresponds to a path through the graph. I have written software to build a splice graph from a collection of short reads aligned to a reference genome. This software incorporates variants observed relative to the reference genome as additional paths. I have also written tools to manipulate and traverse such graphs. An application of this graph is correction of noisy full-length RNA transcripts. Such a transcript may be aligned to paths through the graph in order to identify its original sequence.
Jake Bogerd is a senior from Durham, NC majoring in computer science and mathematics. He works part-time in Research Triangle Park as an intern on the analytics team at Interactive Intelligence. After graduation, he will continue this work full-time as a software engineer.
This talk was given by undergraduate Jake Bogerd during the 10th Annual Computer Science Undergraduate Research Symposium in 2016. Jake‘s research was supervised by Dr. Jan Prins.
“A Method for Construction of a Splice Graph from RNA Sequence Data”
Although genetic information is stored in DNA, the RNA transcripts obtained from DNA indicate which genes are actually active in any given cell. Modern high-throughput sequencing techniques allow accurate sequencing of short RNA fragments, which may then be aligned to a reference genome. These alignments can be summarized by constructing a “splice graph”, in which nodes represent genomic coordinates and edges represent sequences that are retained or spliced out of observed transcripts. Each full-length transcript corresponds to a path through the graph. I have written software to build a splice graph from a collection of short reads aligned to a reference genome. This software incorporates variants observed relative to the reference genome as additional paths. I have also written tools to manipulate and traverse such graphs. An application of this graph is correction of noisy full-length RNA transcripts. Such a transcript may be aligned to paths through the graph in order to identify its original sequence.
Jake Bogerd is a senior from Durham, NC majoring in computer science and mathematics. He works part-time in Research Triangle Park as an intern on the analytics team at Interactive Intelligence. After graduation, he will continue this work full-time as a software engineer.










