发布: 2026年11月05日第16卷第21期 DOI: 10.21769/BioProtoc.5857 浏览次数: 38
评审: Prashanth N SuravajhalaYuhang Wang
Abstract
RNA sequencing (RNA-seq) datasets provide valuable opportunities to investigate gene and transcript abundance across species, tissues, developmental stages, and experimental conditions. However, processing multiple datasets consistently remains challenging because sequencing runs may differ in library layout, read length, sequencing depth, metadata quality, and analytical settings. Existing RNA-seq workflows often depend on dedicated workflow managers, substantial computational infrastructure, or advanced bioinformatics expertise, which may limit their accessibility for routine analyses. We present a modular Bash- and Python-based batch-processing pipeline for estimating gene and isoform abundance and generating expression matrices from multiple RNA-seq runs. Using an SRA RunTable, a reference genome, and its corresponding gene annotation, the workflow automates reference preparation, sequencing-data retrieval, quality assessment, read preprocessing, alignment, abundance estimation, and post-processing. It produces gene-level TPM and FPKM matrices, individual gene- and isoform-level RSEM outputs, quality-control summaries, and integrated MultiQC reports. Configurable computational resources, sample-level status tracking, automatic download retries, selective reprocessing of failed samples, and controlled removal of intermediate files allow interrupted analyses to resume without repeating completed runs while reducing storage requirements. By combining batch processing, transparent configuration, and restartable execution in a lightweight workflow, this protocol provides an accessible approach for standardized RNA-seq abundance estimation.
Key features
• Processes multiple RNA-seq runs in batch, from sequencing-data retrieval and preprocessing to gene- and isoform-level abundance estimation.
• Supports paired-end and single-end libraries with configurable references, preprocessing parameters, and computational resources.
• Enables restartable execution with sample-level tracking, automatic download retries, selective reprocessing, and configurable cleanup of intermediate files.
• Generates gene-level TPM and FPKM matrices, gene- and isoform-level RSEM outputs, quality-control summaries, and aggregated MultiQC reports.
Keywords: TranscriptomicsGraphical overview
Reproducible workflow for standardized processing of public RNA-seq datasets. The pipeline automates data retrieval, preprocessing, expression quantification, quality-control reporting, and result generation from public sequencing data.
Background
Public RNA-seq repositories provide access to extensive datasets generated across species, tissues, developmental stages, and experimental conditions. The NCBI Sequence Read Archive (SRA) enables the reuse of raw sequencing data for transcriptomic analyses beyond the objectives of the original studies [1]. These resources facilitate comparative analyses, hypothesis generation, and data integration without requiring additional sequencing experiments [2]. However, datasets generated by independent studies often differ in sequencing platform, read length, library layout, sequencing depth, metadata quality, and analytical processing, hindering their standardized integration and comparative interpretation. Several computational strategies are available for RNA-seq processing. Alignment-based workflows commonly use splice-aware aligners such as STAR [3] or HISAT2 [4], whereas lightweight approaches rely on pseudoalignment or selective-alignment tools such as kallisto [5] and Salmon [6]. In this pipeline, STAR is a requirement of the architecture rather than an interchangeable preference, because RSEM quantifies directly from the transcriptome-coordinate BAM file that STAR produces via --quantMode TranscriptomeSAM; existing protocols guide users through quality assessment, read preprocessing, alignment, transcript abundance estimation, differential expression analysis, and downstream visualization [7,8]. More comprehensive solutions based on workflow managers, including Nextflow and nf-core/rnaseq, provide scalable and portable processing across computing environments [9,10]. Nevertheless, these approaches may require familiarity with workflow-management systems, substantial computational infrastructure, or advanced bioinformatics expertise. Furthermore, the selection and version of alignment and quantification tools can influence abundance estimates and downstream analyses, particularly for low-abundance genes and transcripts with complex isoform structures [11,12]. This protocol presents a modular batch-processing pipeline for RNA-seq abundance estimation from multiple sequencing runs. The workflow integrates data retrieval, quality assessment, read preprocessing, alignment, transcript abundance estimation, and matrix generation within a transparent and configurable framework. Gene- and isoform-level abundances are estimated using RSEM, which accounts for ambiguously mapped reads and has demonstrated competitive performance relative to alternative RNA-seq quantification methods [11,13]. The pipeline generates standardized TPM and FPKM gene abundance matrices together with individual RSEM abundance files, quality-control summaries, and aggregated MultiQC reports. Transcript-level estimates, reported as individual isoform-level (.isoforms.results) outputs, are deliberately preserved alongside the gene-level files to support downstream gene-level inference through tools such as tximport combined with DESeq2 or edgeR, following Soneson et al. [14], when differential expression analysis is the end goal. This pipeline is executed directly on Linux workstations or servers and does not require a dedicated workflow manager. It incorporates sample-level status tracking, automatic download retries, restartable execution, configurable computational resources, and controlled removal of intermediate files. These features enable reproducible and storage-efficient batch processing of multiple RNA-seq runs while avoiding unnecessary repetition of successfully completed analyses. The workflow can be applied to any organism for which compatible reference genomes, gene annotations, and RNA-seq metadata are available.
Equipment
1. Linux workstation or server with a multicore processor and Bash support
2. Random-access memory (RAM), at least 32 GB recommended
3. Local or network storage with sufficient capacity for SRA, FASTQ, BAM, reference, and result files
4. Stable broadband internet connection for downloading sequencing data, reference files, and software dependencies
Note: The workflow is hardware-independent and can be executed on different workstation or server configurations. Therefore, manufacturer and model information are not applicable.
Software and datasets
Software
1. SRA Toolkit v3.4.1 [15]; public domain; released 2026-04-19
2. FastQC v0.12.1 [16]; GPL ≥3; released 2023-03-04
3. MultiQC v1.14 [17]; GPLv3; released 2023-01-08 (requires setuptools <81, pinned in environment.yml)
4. BBMap/BBDuk v39.81 [18]; BSD-3-Clause-LBNL; released 2026-03-31
5. STAR v2.7.10a [3]; GPLv3; released 2022-01-15
6. RSEM v1.3.3 [13]; GPL-3.0-or-later; released 2025-10-14 (the tool's own --version banner reports v1.3.1)
7. Python v3.11.16 [19]; Python-2.0; released 2026-09-02 (exact pin; not "or later")
8. openpyxl v3.1.5 [20]; MIT; released 2026-09-04 (exact pin)
9. GNU Wget v1.25.0 [21]; GPL-3.0-or-later; released 2026-02-27 (exact version used in validation)
10. Conda v26.5.3 [22]; BSD-3-Clause; released 2026-06-16
11. Git ≥ 2.30 [23] (used to clone the repository; not Conda-managed)
12. Bash ≥ 4.4 [24] (uses mapfile, printf -v, and associative arrays)
13. curl or GNU Wget [25] (either satisfies reference/RunTable downloads)
14. flock (util-linux) [26], optional (enables the run lock; without it, run.sh starts unlocked)
Note: All Conda-managed versions above are exact pins in environment.yml (not minimum or "or later" versions). Release dates are those of the exact package build installed from its channel, which is the date that matters for reproduction. A frozen, fully explicit package list (environment.lock.txt, generated with conda list --explicit) is archived with the Zenodo DOI for exact reproduction of the validated environment.
Required input files
• SRA RunTable (CSV or XLSX)
• Reference genome (FASTA)
• Gene annotation (GTF)
Note: This workflow supports standard paired-end and single-end RNA-seq libraries, extracted with --split-3 during FASTQ conversion. Mate-pair libraries are a long-insert genomic preparation rather than an RNA-seq library type and are therefore out of scope. Processing mate-pair data would require different STAR alignment settings, including mate orientation and --alignMatesGapMax, and is not supported by the current pipeline version.
Procedure
登录/注册后免费查看全文
文章信息
稿件历史记录
提交日期: Aug 6, 2026
接收日期: Sep 17, 2026
在线发布日期: Oct 10, 2026
出版日期: Nov 5, 2026
版权信息
© 2026 The Author(s); This is an open access article under the CC BY license (https://creativecommons.org/licenses/by/4.0/).
如何引用
Leiton, J. S. Z., López, K. E., Riascos-España, A. F., Gonzalez, C. E. S., García, C. A. B. and Velasquez-Vasconez, P. A. (2026). A Batch-Processing Pipeline for RNA-seq Gene Abundance Estimation. Bio-protocol 16(21): e5857. DOI: 10.21769/BioProtoc.5857.
您对这篇实验方案有问题吗?
在此处发布您的问题,我们将邀请本文作者来回答。同时,我们会将您的问题发布到Bio-protocol Exchange,以便寻求社区成员的帮助。
Share
Bluesky
X
Copy link

