Bulk RNA-sequencing (whole transcriptome)
| Definition: High-throughput sequencing of RNA transcribed by an individual or population of a given organism, or a community of species. |
| Approach: filtering, RNA extraction, sequencing, bioinformatics |
| Context: in situ, culturing, incubations |
| Spatial scale: L |
| Temporal scale: hours, days, weeks, seasons, events of interest |
| Units: gene expression, e.g. TPM |
| Community captured: size-fractioned, e.g. 0.2 - 250 µm, cultured or target species (both prokaryotic and eukaryotic) |
| Co-measurements: Other measurements required for interpretation of in-situ samples e.g., temperature, salinity, nutrients, physiological data |
Method Overview
(Meta)transcriptomics sequences the pool of RNA recovered from a single-species culture up to a whole microbial assemblage. In the latter case it can provide a semi-quantitative snapshot of which genes were being transcribed by which organisms at the moment of sampling. Where (meta)genomics describes metabolic potential (i.e. the genes that are present), (meta)transcriptomics describes metabolic expression (i.e. the genes that are being transcribed), and is therefore the molecular measurement most often used as a proxy for in situ physiological state and for the activity of specific biogeochemical pathways.
The method is sensitive to how the sample is handled. Prokaryotic mRNA has a half-life of roughly 5 minutes, and as little as ~2.4 minutes in marine cyanobacteria.[1] Protocol choices at every step during the generation of bulk RNA-Seq data such as filtration, preservation, enrichment (e.g. for mRNA in poly(A) selection to target microeukaryotic material), sequencing depth, assembly and annotation, affect the final expression matrix[2]
Sample collection and filtration
Seawater can be collected by Niskin bottles, underway pump, manually, or in situ pump, and biomass is concentrated onto filters. Culture samples can be taken directly from the culture vessel (potentially with upconcentration). Because the transcriptome degrades and responds to handling stress within minutes, the interval between sampling and freezing should be minimised[2].
- Volume. Can be several mL from cultures, 1.5-3.5 L in productive coastal waters, and ≥10 L in oligotrophic offshore waters for microeukaryote work; 1–10 L is common for prokaryote-targeted sampling on 0.22 µm Sterivex cartridges. Volume is constrained by filter clogging and by the requirement to complete filtration quickly. More volume can be filtered and later subsampled.
- Size fractionation. Serial or stacked filters partition the assemblage, e.g. 0.2–3 µm (prokaryote-enriched), 3–20 µm and 20–250 µm (nano- and microplankton), or a single 0.8–250 µm protistan fraction. Fractionation aids interpretation but introduces cell breakage and incomplete separation, and the chosen cut-offs must accompany any downstream comparison.
- Preservation, extraction and quality control. Filters or filtrate can be flash-frozen in liquid nitrogen and stored at −80 °C; RNA-stabilising buffers (e.g. RNAlater) can be added to safeguard RNA integrity[2]
RNA extraction & library preparation
Internal standards: Adding a known number of synthetic RNA molecules at the start of extraction converts an otherwise compositional (relative) dataset into an absolute one. Standards may be commercial spike-in sets (ERCC, ArrayControl, Sequins) or in vitro transcripts from custom plasmids, and are typically dosed at ~1% of the expected mRNA pool.[2][3] The ratio of standard molecules added to standard reads recovered yields transcript inventories per litre of seawater (see Common calculations), and also diagnoses extraction and library-prep losses. Quantitative metatranscriptomics of this kind is the form most directly comparable to rate measurements and to model state variables.[4] RNA extraction Total RNA is extracted from the sample often using commercial kits (e.g. Qiagen RNeasy, Invitrogen TRIzol, ToTALLY RNA), usually with bead-beating in silica or zirconia beads to break silicified, armoured or thick-walled cells, followed by DNase treatment to remove co-extracted DNA. Yield and integrity are checked by fluorometry (Qubit) and capillary electrophoresis (Bioanalyzer/TapeStation)[2].
For eukaryotes, ribosomal RNA dominates total RNA (usually >80%), so libraries are enriched for mRNA by one of two routes:
- Poly-A selection captures polyadenylated eukaryotic mRNA on oligo-dT beads. It delivers the most protein-coding reads per sequencing dollar, but excludes bacteria and archaea, discriminates against plastid and mitochondrial transcripts, and biases against transcripts with short or absent poly-A tails. Standard kits (e.g. TruSeq Stranded mRNA) need 0.1–1 µg total RNA; low-input chemistries (e.g. SMART-Seq v4) work from ~250 pg via linear amplification, at the cost of amplification bias.[2]
- rRNA depletion (Ribo-Zero, riboPOOLs and equivalents) removes rRNA by probe hybridisation and retains prokaryotic, eukaryotic and organellar mRNA together. It is the choice for whole-community and prokaryote-focused work and for organelle transcripts, but leaves a larger residual rRNA fraction and so requires greater sequencing depth per unit of usable signal.[2]
Residual rRNA is removed in silico after sequencing with SortMeRNA, BBDuk or riboPicker. Libraries are further prepared according to the sequencing platform.
Sequencing
Short-read Illumina sequencing (NextSeq, NovaSeq; typically 2×150 to 2×250 bp) is the most popular platform.
Bioinformatic processing
- Quality control. Adapter and quality trimming (fastp, Trimmomatic, cutadapt) with inspection via FastQC/MultiQC; in silico rRNA removal.
- Direct mapping or de novo assembly. Reads may be mapped directly to a reference genome or database (e.g. MMETSP, EukProt, MarFERReT, OM-RGC) or assembled de novo (Trinity, rnaSPAdes, MEGAHIT). rnaSPAdes and Trinity perform well on marine microeukaryote communities. Multi-assembler approaches or assemblies from multiple samples can be followed by clustering at 95-100% identity (CD-HIT, MMseqs2)[2].
- Quantification. Trimmed reads are mapped back to the assembly with Bowtie2/BWA or quasi-mapped with Salmon/kallisto to produce per-contig counts.
- Protein prediction and annotation. Open reading frames are called (TransDecoder, GeneMarkS-T; Prodigal for prokaryote-dominated assemblies), then annotated functionally against KEGG/KOfam, eggNOG, Pfam (HMMER), KOG and Gene Ontology, and taxonomically by DIAMOND/MMseqs2 alignment (with a last-common-ancestor algorithm)[2].
- Aggregation and statistics. Quantified matrics are integrated with environmental metadata, and annotation information to tackle questions of interest.
Workflow examples: Reproducible pipelines include eukrhythmic, SqueezeMeta and nf-core/metatdenovo,[5][6] and require high-performance computing resources for large datasets.
Output & normalisation
Sequencing produces compositional data: read counts describe proportions of a fixed-size library. Reads should be normalised:
- Within-sample relative expression (TPM, RPKM/FPKM) — comparable across genes within a sample, and the pragmatic choice for large spatial or temporal surveys, but composition-dependent and without a statistical test of differential expression.[2]
- Model-based differential expression (DESeq2, edgeR) — negative-binomial generalised linear models with FDR control, appropriate for replicated designs and ideally applied after binning to taxon-specific transcript sets.[2]
- Absolute estimates (transcripts L-1, transcripts cell-1) — requires internal standards(and logging of processed volumes, optionally with cell counts or a paired metagenome), and is the form best suited to comparison with rates and with model fluxes.[4][3][7]
Scale of measurement
Next to culture-based approaches, in-situ sampling of bulk RNA-Seq depends on the study design. A single sample is a point measurement in space: one sample integrates the litres of seawater it collected (typically 1–10 L, more offshore), at one depth and station. These can then be aggregated across the research cruise, mooring or autonomous platform, individual snapshots resolve diel cycles, bloom development, mesoscale features and seasonal transitions; global-scale synthesis is possible if protocols are harmonised.[8][9]
Data generated
- Raw paired-end sequence reads (FASTQ), with library and sample metadata.
- Assembled contigs or a mapped reference catalogue (FASTA), and predicted proteins.
- A count matrix of contigs/genes × samples, plus taxonomic and functional annotation tables.
- Derived products: normalised expression matrices (TPM), differential-expression tables, transcript inventories per litre, ordinations, co-expression networks, and pathway- or taxon-level expression profiles.
- Where internal standards were used, standard recovery statistics documenting extraction and library efficiency.
- Environmental or study metadata
Repositories & databases
Raw and processed data
- NCBI Sequence Read Archive (SRA) and ENA — raw reads; the expected archive for publication.
- Zenodo and GitHub for processed data ( assemblies, count matrices, metadata, etc) and analysis code[2].
Reference libraries for annotation
- MMETSP (Marine Microbial Eukaryote Transcriptome Sequencing Project): 678 transcriptomes from 405 microbial eukaryote strains.[10]
- MarFERReT: version-controlled reference library of marine microbial eukaryote functional genes.[11]
- EukProt, EukZoo and PhyloDB: complementary eukaryote protein references; MarRef and OM-RGC for prokaryotes.
- Ocean Gene Atlas: online query of Tara Oceans gene abundance and expression biogeography.[12]
- North Pacific Eukaryotic Gene Catalog: assembled and annotated metatranscriptomes for the North Pacific.[13]
- Functional references: KEGG/KOfam, eggNOG, Pfam, KOG, Gene Ontology.
Limitations
Sampling and handling
- mRNA half-lives of minutes mean the measurement is easily perturbed. Time from collection to preservation, bottle confinement, light and temperature change, and filtration pressure all induce stress responses that are indistinguishable from in situ signal unless handling is standardised.[1][2]
- Strong diel periodicity in transcription means samples are only comparable when time of day is matched or explicitly modelled[9][14]
- Size fractionation is imperfect: cells break, fractions overlap, and filters clog, biasing which taxa are represented.
Library chemistry
- Poly-A selection excludes prokaryotes and discriminates against organellar and non-polyadenylated transcripts; rRNA depletion retains more rRNA and needs deeper sequencing. Neither is bias-free, and libraries prepared by the two routes are not directly comparable.[2]
Quantification and statistics
- Without internal standards the data are compositional: an apparent increase in one transcript can reflect a decrease elsewhere.[3][7]
- Sequencing captures a minute fraction of the mRNA pool, so many functionally important genes fall below detection; roughly half the functional categories in early quantitative work had too few counts for robust comparison.[4]
- Limited biological replication constrains formal differential-expression testing.[2]
Reference and assembly
- Reference libraries under-represent rare lineages, deep-sea taxa and uncultured groups and often contain contamination; a large share of contigs and predicted proteins remain taxonomically or functionally unannotated.[2]
- De novo assembly of mixed communities generates chimeric and spurious contigs where close relatives co-occur, and results depend on assembler and k-mer choice.[2]
- Choosing a taxonomic resolution is a trade-off: species-level assignment is confident for fewer sequences, while class-level aggregation can mask opposing physiological responses within a group. The use of reference databases suffer from different biases.[2][15]
Interpretation
- Transcript abundance is not protein abundance and not a rate.
- Expression change can arise from a shift in community composition rather than in per-cell regulation[8].
- Converting expression into a flux requires independent calibration (rate measurements, physiological data, or a model).
Example studies
https://doi.org/10.1073/pnas.0708897105
https://doi.org/10.1111/j.1462-2920.2008.01863.x
https://doi.org/10.1073/pnas.1118408109
https://doi.org/10.1073/pnas.1421993112
https://doi.org/10.1038/s41467-017-02342-1
https://doi.org/10.1016/j.cell.2019.10.014
References
- ↑ 1.0 1.1 Moran, M. A., Satinsky, B., Gifford, S. M., Luo, H., Rivers, A., Chan, L.-K., Meng, J., Durham, B. P., Shen, C., Varaljay, V. A., Smith, C. B., Yager, P. L., & Hopkinson, B. M. (2013). Sizing up metatranscriptomics. The ISME Journal, 7(2), 237–243. doi:10.1038/ismej.2012.94
- ↑ 2.00 2.01 2.02 2.03 2.04 2.05 2.06 2.07 2.08 2.09 2.10 2.11 2.12 2.13 2.14 2.15 2.16 2.17 Cohen, N. R., Alexander, H., Krinos, A. I., Hu, S. K., & Lampe, R. H. (2022). Marine microeukaryote metatranscriptomics: sample processing and bioinformatic workflow recommendations for ecological applications. Frontiers in Marine Science, 9, 867007. doi:10.3389/fmars.2022.867007
- ↑ 3.0 3.1 3.2 Satinsky, B. M., Gifford, S. M., Crump, B. C., & Moran, M. A. (2013). Use of internal standards for quantitative metatranscriptome and metagenome analysis. Methods in Enzymology, 531, 237–250. doi:10.1016/B978-0-12-407863-5.00012-5
- ↑ 4.0 4.1 4.2 Gifford, S. M., Sharma, S., Rinta-Kanto, J. M., & Moran, M. A. (2011). Quantitative analysis of a deeply sequenced marine microbial metatranscriptome. The ISME Journal, 5(3), 461–472. doi:10.1038/ismej.2010.141
- ↑ Di Leo, et al. (2025). The Nextflow nf-core/metatdenovo pipeline for reproducible annotation of metatranscriptomes, and more. PeerJ, 13, e20328. doi:10.7717/peerj.20328
- ↑ Krinos, A. I., Cohen, N. R., Follows, M. J., & Alexander, H. (2023). Reverse engineering environmental metatranscriptomes clarifies best practices for eukaryotic assembly. BMC Bioinformatics, 24, 74. doi:10.1186/s12859-022-05121-y
- ↑ 7.0 7.1 Perneel, M., Alexander, H., Hablützel, P. I., & Maere, S. (2025). A case for absolute gene expression estimates in microbiome studies using metatranscriptomics. The ISME Journal, 19(1), wraf188. doi:10.1093/ismejo/wraf188
- ↑ 8.0 8.1 Salazar, G., Paoli, L., Alberti, A., et al. (2019). Gene expression changes and community turnover differentially shape the global ocean metatranscriptome. Cell, 179(5), 1068–1083.e21. doi:10.1016/j.cell.2019.10.014
- ↑ 9.0 9.1 Peoples, L. M., Eppley, J. M., Barone, B., Hobson, B. W., Karl, D. M., Kieft, B., Marin, R. III, Preston, C. M., Romano, A. E., Ryan, J. P., Scholin, C. A., Wilson, S. T., Zhang, Y., Church, M. J., & DeLong, E. F. (2026). Diel and eddy driven changes in microbial gene expression and biogeochemistry in the oceanic chlorophyll maximum. Nature Communications, 17, 3636.
- ↑ Keeling, P. J., et al. (2014). The Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP): illuminating the functional diversity of eukaryotic life in the oceans through transcriptome sequencing. PLoS Biology, 12(6), e1001889. doi:10.1371/journal.pbio.1001889
- ↑ Groussman, R. D., Blaskowski, S., Coesel, S. N., & Armbrust, E. V. (2023). MarFERReT, an open-source, version-controlled reference library of marine microbial eukaryote functional genes. Scientific Data, 10, 926. doi:10.1038/s41597-023-02842-4
- ↑ Vernette, C., Lecubin, J., Sánchez, P., Tara Oceans Coordinators, Sunagawa, S., Delmont, T. O., Acinas, S. G., Pelletier, E., Hingamp, P., & Lescot, M. (2022). The Ocean Gene Atlas v2.0: online exploration of the biogeography and phylogeny of plankton genes. Nucleic Acids Research, 50(W1), W516–W526. doi:10.1093/nar/gkac420
- ↑ Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. Scientific Data, 11, 1161. doi:10.1038/s41597-024-04005-5
- ↑ Ottesen, E. A., Young, C. R., Gifford, S. M., Eppley, J. M., Marin, R. III, Schuster, S. C., Scholin, C. A., & DeLong, E. F. (2014). Multispecies diel transcriptional oscillations in open ocean heterotrophic bacterial assemblages. Science, 345(6193), 207–212. doi:10.1126/science.1252476
- ↑ Krinos, A.I., Mars Brisbin, M., Hu, S.K. et al. Missing microbial eukaryotes and misleading meta-omic conclusions. Nat Commun 15, 9873 (2024). https://doi.org/10.1038/s41467-024-52212-w