Building an Oxford Nanopore Pipeline for RNA-Seq Data Analysis

Developing an end-to-end long-read transcriptomics pipeline using Oxford Nanopore sequencing: basecalling, splice-aware alignment, isoform discovery, and differential expression.

Kushan Manahara

May 20, 2024 · 4 min read

00
Building an Oxford Nanopore Pipeline for RNA-Seq Data Analysis

Our final year project at the University of Peradeniya was titled 'Developing an Oxford Nanopore Technology Based Pipeline for RNA-Seq Data Analysis'. When we first took on the topic, bioinformatics was an entirely new domain for our team. We had strong foundations in computer systems, algorithms, and high-performance computing, but biology brought a fascinating new level of complexity.

Traditional short-read sequencing (such as Illumina) produces fragments of 150–300 base pairs. While highly accurate, assembling these tiny jigsaw pieces to identify complex full-length RNA transcript isoforms is computationally ambiguous. Oxford Nanopore Technologies (ONT) revolutionizes transcriptomics by reading entire RNA or cDNA molecules in single continuous passes, from 5' end to poly-A tail.

The Computational Pipeline Architecture

Because native ONT long-read RNA sequencing has unique error profiles and massive raw data volumes (gigabytes of raw electrical signal data), standard short-read pipelines fail. We researched, engineered, and benchmarked an end-to-end computational pipeline:

  • GPU-Accelerated Basecalling (Dorado): Translating raw ionic current fluctuations (POD5/FAST5 files) into raw nucleotide sequence reads using deep neural network basecallers on NVIDIA GPUs.
  • Quality Control & Filtering (NanoPlot & Chopper): Evaluating per-read Phred quality scores (Q-score distributions) and length metrics, filtering out truncated adapters and chimeric artifacts.
  • Splice-Aware Sequence Alignment (Minimap2): Aligning long cDNA/direct-RNA reads to the reference genome using minimap2 -ax splice, specifically tuned to account for non-canonical splice junctions and variable intron sizes.
  • Transcript Quantification & Isoform Discovery (StringTie2): Reconstructing complete transcript models, identifying novel alternative splicing events, and calculating gene/transcript abundance metrics (TPM).
  • Differential Expression (DESeq2 in R): Statistical modeling of negative binomial distributions across biological replicates to detect significantly up- and down-regulated genes.

Key Systems Challenges Solved

From a systems engineering perspective, bioinformatics is a stress test for hardware. We encountered two major bottlenecks:

1. I/O Bottlenecks: Processing hundreds of thousands of individual raw signal files overwhelmed standard hard drives. Migrating raw working scratch disks to NVMe arrays and batching reads drastically reduced pipeline idle time.

2. GPU Memory Allocation: Basecalling with high-accuracy models required careful chunking and memory pin management to avoid CUDA out-of-memory errors during long sequencing runs.

Acknowledgments

I had the privilege of working on this research alongside my project partners Chiran Govinna and Tharindu Dhananjaya. We embraced every hurdle, solved stubborn bugs together, and successfully defended our final thesis.

Our deepest gratitude goes to our supervisor Dr. Asitha Bandaranayake for believing in our engineering capabilities and providing tireless guidance. Sincere thanks to Prof. Pradeepa C.G. Bandaranayake for her invaluable domain expertise in molecular biology and bioinformatics, and to Dr. Bhagya Chandrasekara whose day-to-day research assistance was indispensable in navigating biological datasets.

Written by

Kushan Manahara

Responses (0)

Verified name, role, and email required before posting.

No responses yet

Be the first to share your thoughts, benchmarks, or feedback above.