• Home page
  • Resources
  • BLOG
  • How Does ONT Long-Read Transcriptomics Reveal Elusive Disease Regulatory Mechanisms?

How Does ONT Long-Read Transcriptomics Reveal Elusive Disease Regulatory Mechanisms?


Release time:2026-08-29 09:26:50


1、Introduction: A New Era of Disease Transcriptomics Research

The transcriptome serves as a critical bridge connecting the genome and phenome, occupying a core position in deciphering gene regulatory networks, mining disease biomarkers, and elucidating pathogenic molecular mechanisms. Over the past decade, Illumina short-read sequencing has driven the rapid development of transcriptomics, enabling routine genome-wide quantification of gene expression.

However, accumulating scientific evidence indicates that changes in gene expression levels alone are insufficient to explain the high complexity of human diseases. Approximately 95% of human multi-exon genes undergo alternative splicing, generating diverse transcript isoforms that encode proteins with distinct or even opposing biological functions. Isoform switching and differential expression are proven to be pivotal regulatory events underlying the occurrence and progression of tumors, neuropsychiatric disorders, cardiovascular diseases, immune disorders, and multiple other illnesses.

Constrained by short read lengths, traditional second-generation sequencing fails to span multiple exons of complete transcripts, leaving a large number of functional isoforms unrecognized and turning them into "dark matter" in disease transcriptional research. Oxford Nanopore Technologies (ONT) full-length transcriptome sequencing eliminates the need for RNA fragmentation, enabling direct capture of complete RNA sequences from the 5' end to the 3' poly(A) tail. This cutting-edge technology allows researchers to fully visualize the complete structural information of individual transcripts, opening a new dimension for in-depth exploration of complex disease transcriptional regulatory mechanisms.

 

2、Current Challenges: Bottlenecks of Short-Read Sequencing in Disease Research

While second-generation short-read sequencing has laid a solid foundation for transcriptomic research, its inherent technical limitations have gradually become a major bottleneck restricting precise exploration of disease transcriptional regulation, resulting in systematic deviations and omissions in transcript identification, quantification, and structural analysis:

Incomplete transcript identification: Short-read sequencing requires RNA fragmentation into 200–500 bp segments followed by computational splicing. For genes with multiple highly similar isoforms, short fragments cannot be accurately traced to their original transcripts, leading to frequent splicing errors and widespread isoform missing detection. For instance, in acute myeloid leukemia research, over 31.16% of unannotated isoforms discovered by ONT sequencing were completely undetectable via short-read platforms.

Inaccurate transcript quantification: Short reads cannot be uniquely mapped to specific isoforms, limiting traditional transcriptomic analysis to gene-level quantification. Notably, different isoforms derived from the same gene may exert completely opposite biological functions (e.g., oncogenic vs. tumor-suppressive effects). Mixing the expression signals of functionally distinct isoforms inevitably leads to one-sided or even misleading research conclusions.

Limited identification of structural variations: Key transcriptional structural variations, including alternative splicing, gene fusion, and alternative polyadenylation (APA), rely on complete sequence spanning of splicing junctions and fusion breakpoints. Restricted by read length, short-read sequencing can only capture partial fragment information, making it impossible to fully and accurately characterize these critical disease-related structural variations.

Loss of RNA modification information: The library construction process of short-read sequencing requires reverse transcription and PCR amplification, which completely erases post-transcriptional RNA modifications such as m6A and 5mC. These epigenetic modifications are core regulatory elements of transcriptional post-modification and are closely associated with the pathogenesis of neurodegenerative diseases, cancers, and other major disorders.

Terminal sequence coverage bias: Random RNA fragmentation causes uneven coverage of transcript 5' and 3' terminals, resulting in inaccurate terminal sequence information and hindering the precise identification of transcript initiation and termination sites.

Collectively, these technical limitations confine short-read transcriptomic research to coarse-grained gene-level analysis, failing to support fine-grained transcript-level regulatory exploration — the core key to unraveling the complex molecular mechanisms of human diseases.

 

3、Technological Breakthroughs: Advantages and Scientific Value of ONT Full-Length Transcriptome

3.1 Technical Principle: From Fragment Splicing to Full-Spectrum Transcript Reading

ONT full-length transcriptome sequencing is based on single-molecule nanopore sensing technology. It eliminates the mandatory fragmentation step of traditional RNA sequencing. Either full-length cDNA reverse-transcribed from total RNA or native intact RNA molecules can be directly sequenced through nanopores. The platform supports ultra-long reads ranging from 10–20 kb (maximum exceeding 80 kb), which can cover almost all full-length transcripts of eukaryotic genes.

The fundamental technological difference is revolutionary: short-read sequencing follows a "fragment first, splice later" logic that relies on algorithmic inference, while ONT sequencing adopts a "direct full-length reading" mode that captures real transcript structures without secondary splicing.

 

6fedc62d-8a44-4c75-ac58-dc1341b28d9e_看图王.jpg

 

Figure 1 Comparison of short-read and ONT long-read transcriptome sequencing

(Source: https://nanoporetech.com/api/assets/f/196663/x/6caaeff1e8/bulk-transcriptomics-getting-started-guide.pdf)

As shown in Figure 1, short reads cannot fully cover single exons and require complex algorithmic splicing with high error rates. In contrast, ONT long reads directly capture complete transcript sequences from start to end, accurately restoring the authentic structural characteristics of transcripts.

3.2 Core Technical Advantages: Solving Key Pain Points in Transcriptomic Research

Accurate identification of full-length transcripts: ONT sequencing directly acquires complete transcript sequences containing 5'UTR, CDS, 3'UTR, and poly(A) tails. It achieves unambiguous identification of individual isoforms, fundamentally avoiding splicing errors and chimeric transcript artifacts prevalent in short-read data.

Precise isoform-level quantification: Each ONT read corresponds to an independent complete transcript, enabling direct and accurate quantification of expression abundance at the isoform level. This provides reliable data support for studying isoform switching events closely linked to disease occurrence and progression. Benchmark studies have verified that long-read sequencing outperforms short-read platforms significantly in identifying dominant functional isoforms.

Comprehensive detection of complex structural variations: Ultra-long reads effortlessly span multiple exons, variable splicing junctions, and distant gene fusion breakpoints, enabling comprehensive and accurate identification of alternative splicing, fusion genes, APA, and other structural variations that are undetectable or inaccurately detected by short-read sequencing.

Retention of complete RNA modification information: The ONT direct RNA sequencing (DRS) mode sequences native RNA molecules without amplification, preserving natural base modification signatures including m6A, pseU, inosine, m5C, and 2’OMe. It adds a new dimension for in-depth study of post-transcriptional regulatory mechanisms in diseases.

Elimination of PCR and GC bias: Based on electrical signal detection, the technology requires no polymerase amplification, ensuring true and unbiased detection and quantification of low-abundance transcripts and genes with extreme GC content.

3.3 Scientific and Application Value

The core innovation of ONT full-length transcriptome technology is elevating transcriptomic research from conventional gene-level analysis to high-precision transcript isoform-level analysis. In disease research scenarios, it enables researchers to break through the limitations of traditional research logic — shifting from the superficial conclusion of "abnormal gene expression" to in-depth exploration of "which isoform is dysregulated, how structural variations affect protein function, and how isoform switching drives disease progression".

This technological upgrade is particularly critical for precision disease research. For example, many tumor suppressor genes and oncogenes have multiple isoforms with completely opposite functions. Gene-level mixed expression analysis cannot distinguish pathogenic functional isoforms, while ONT transcript-level resolution accurately screens disease-driving isoforms, providing precise targets for disease mechanism research, biomarker discovery, and targeted therapy development.

3.4 Full Technical Workflow

ONT full-length transcriptome research covers a closed-loop workflow from sample processing to biological interpretation, realizing standardized and high-precision omics analysis:

RNA Extraction & Quality Control → Full-Length Library Construction → PromethION Platform Sequencing → Bioinformatic Analysis (full-length transcript identification & quantification, alternative splicing analysis, fusion gene detection, APA analysis, novel transcript mining, isoform differential expression analysis, RNA modification identification, multi-omics integration) → Biological Interpretation (disease biomarker screening, drug target mining, molecular subtyping, prognostic evaluation)

 

5240507b-e325-4856-a5ae-326fdb892ec3_看图王.jpg

 

Figure 2 Workflow of ONT full-length transcriptome sequencing and data analysis

(Source: https://academic.oup.com/bib/article/27/1/bbag063/8482848?login=false

 

4、Research Case: ONT Full-Length Transcriptome Decodes Colorectal Cancer Transcriptional Mechanisms

Paper Title: Long-read sequencing reveals the landscape of aberrant alternative splicing and novel therapeutic target in colorectal cancer

Journal: Genome Medicine (IF=12.3)

Publication Date: September 21, 2023

DOI: 10.1186/s13073-023-01226-y

Alternative splicing dysregulation is a core hallmark of colorectal cancer (CRC) tumorigenesis and progression, yet the full landscape of CRC-associated splicing variations cannot be fully resolved by traditional short-read sequencing. In this landmark study, researchers adopted a combined strategy of ONT long-read sequencing and Illumina short-read sequencing to systematically profile the transcriptome of 78 CRC tumor tissues and 10 adjacent normal tissues.

The study identified a total of 90,703 transcripts, among which more than 62% were unannotated novel transcripts, successfully constructing the most comprehensive landscape of aberrant alternative splicing in CRC to date. Further functional verification confirmed that the newly discovered TIMP1 Δ4-5 splicing isoform acts as a tumor suppressor that significantly inhibits CRC cell proliferation and metastasis, while the full-length TIMP1 isoform plays an oncogenic role. Mechanistically, the splicing factor SRSF1 drives CRC progression by maintaining the inclusion of TIMP1 exons 4–5. Based on this finding, the study developed a CRISPR/dCasRx splicing editing strategy to induce TIMP1 exon 4–5 skipping, which effectively inhibited tumor growth in vivo.

Core Value of ONT Technology in This Research

The breakthrough achievements of this study are entirely dependent on the unique advantages of ONT full-length sequencing. Short-read sequencing, with only 100–150 bp read length, cannot cover complete variable splicing events, making it impossible to distinguish real functional isoforms from algorithmic splicing artifacts. In contrast, ONT technology directly captures complete single-molecule transcript sequences without fragmentation and splicing inference, providing solid sequence evidence for every novel isoform.

This enables researchers to accurately screen disease-specific functional isoforms from massive transcript data and systematically characterize the dysregulated splicing network in CRC. Beyond CRC mechanism exploration, this study establishes a universal methodological paradigm for tumor transcriptome research, fully proving that isoform-level transcriptional diversity is a key molecular basis for tumor occurrence. It strongly demonstrates the irreplaceable value of ONT full-length transcriptome in mining novel disease biomarkers and therapeutic targets.

 

da341623-b5c7-4f51-9c0a-3b722e6e7d47_看图王.jpg

 

Figure 3 Novel transcript identification in CRC via ONT long-read sequencing

(Source: https://rd.springer.com/article/10.1186/s13073-023-01226-y)

 

5、Sailgene Professional ONT Full-Length Transcriptome Solutions

Sailgene is a professional biotechnology company focusing on long-read sequencing, NGS sequencing, and integrated bioinformatics solutions. We provide comprehensive multi-omics services covering genomics, transcriptomics, microbiomics, and epigenomics, serving scientific research and industrial applications in biomedicine, agriculture, animal husbandry, and environmental science.

In the field of ONT full-length transcriptome sequencing, Sailgene has accumulated rich practical experience, covering hundreds of species including humans, mice, primates, livestock, and plants. We have completed numerous high-complexity disease sample sequencing projects and assisted customers in publishing a large number of high-level research papers.

Our one-stop ONT full-length transcriptome services include:

  1. Full-process technical support: Customized experimental design, professional sample processing, diverse library construction modes, and high-throughput sequencing based on ONT PromethION platform, delivering standardized and high-quality raw data.

  2. Professional in-depth bioinformatic analysis: Covering full-length transcript identification and quantification, alternative splicing profiling, fusion gene detection, APA analysis, novel transcript mining, isoform differential expression analysis, and RNA modification identification.

  3. Multi-omics integrated analysis: Combining transcriptome data with genomics, epigenomics, and proteomics data to achieve multi-dimensional systematic interpretation of disease molecular mechanisms.

  4. Customized solution development: Tailoring personalized sequencing and analysis schemes according to different disease models and research objectives to meet differentiated scientific research needs.

With rigorous quality control standards and professional technical capabilities, Sailgene helps researchers break through the limitations of traditional transcriptomics and accurately mine core disease-driving transcripts and regulatory mechanisms.

 

6、FAQ: Common Technical Questions & Answers

Q1: Why is long-read sequencing necessary for transcriptomic research?

Traditional short-read sequencing with 100–150 bp read length cannot cover complete multi-exon transcript structures. For genes generating multiple highly similar isoforms, short fragments cannot be accurately mapped to specific isoforms, and splicing algorithms cannot distinguish real splicing events from computational errors, resulting in massive missing and misidentification of functional isoforms. In contrast, ONT long-read sequencing directly reads complete transcript sequences, enabling unambiguous identification of authentic isoform structures and achieving accurate isoform-level expression quantification, which is essential for in-depth disease transcriptional research.

Q2: What is the recommended data volume for ONT full-length transcriptome sequencing?

For conventional transcriptomic research, a data volume of 6G is generally sufficient to meet basic transcript identification and quantification requirements. For research on complex species, low-abundance functional transcripts, and in-depth alternative splicing mechanism exploration, it is recommended to increase the data volume to 10G or higher to ensure comprehensive coverage of rare isoforms and subtle splicing variation information.

 

Contact Us

If you are interested in our long-read sequencing services or potential collaboration, please contact us. Our team is ready to support your research with tailored solutions. We also welcome feedback from users to help us improve our services.

Contact Us
%{tishi_zhanwei}%

Contact Us

E-mail:service@sailgene.com

Inquiry

Tel:16172237544

Email:service@sailgene.com

中企跨境-全域组件 制作前进入CSS配置样式

在线客服添加返回顶部

右侧在线客服样式 1,2,3 1

图片alt标题设置: SAILGENE TECHNOLOGY INC.

表单验证提示文本: Content cannot be empty!

循环体没有内容时: Sorry,no matching items were found.

CSS / JS 文件放置地

Welcome to leave an online message, we will contact you promptly

%{tishi_zhanwei}%