Bioinformatics Notes

Transcriptomics and Differential Gene Expression

Lecture slides and notes for Transcriptomics and Differential Gene Expression in Bioinformatics Notes by Md Ahbab. 12 pages.

Document Info: 12 pages ยท PDF

Transcriptomics and Differential Gene Expression, first page preview

Content Preview

CSE 4893 Introduction to Bioinformatics Module 7 CSE 4893: Introduction to Bioinformatics Module 7: Transcriptomics and Differential Gene Expression Abstract These notes cover the statistical core of bulk RNA sequencing analysis. We follow onedataset from library preparation to a ranked results table, passing throughquantification, normalisation, count modelling and multiple testing. The negativebinomial generalised linear model shared by DESeq2 and edgeR is stated in full, togetherwith median of ratios size factors, the trimmed mean of M values, dispersion shrinkage,Wald and likelihood ratio tests, and the Benjamini and Hochberg procedure. Worked numericexamples, pseudocode and compact reference tables re- place long prose. By the end youshould be able to read a differential expression report critically and defend everymodelling choice inside it. Contents 1 Learning Objectives 2 2 From RNA to Counts 2 3 Experimental Design 3 4 Normalisation 3 4.1 Within Sample Units . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 4.2 Between Sample Size Factors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 4.3 Trimmed Mean of M Values

CSE 4893 Introduction to Bioinformatics Module 7 1 Learning Objectives 1. Explain how sequenced fragments become an integer count matrix, and choose between an alignment route and a lightweight quantification route. 2. Distinguish within sample units (CPM, TPM, FPKM) from between sample scaling (median of ratios, TMM), and state when each is appropriate. 3. Write down the negative binomial count model, its mean and variance, and justify it against a Poisson or Gaussian alternative. 4. Build a design matrix for one factor and for two factors, fit the generalised linear model, and test a coefficient with a Wald or likelihood ratio test. 5. Control the false discovery rate with the Benjamini and Hochberg rule and read MA, volcano, PCA and heatmap diagnostics with effect size in mind. 2 From RNA to Counts A library is the physical pool of adapter flanked cDNA fragments that the sequencerreads. Typ- ical steps are RNA extraction and quality scoring, poly(A) selection orribosomal depletion, fragmentation, reverse transcription, adapter ligation, a shortPCR amplification and sequencing. Two choices matter for the statistics later: whether theprotocol is stranded, and whether reads are sin

CSE 4893 Introduction to Bioinformatics Module 7 3 Experimental Design A biological replicate is an independently grown or independently sampled unit. Twolibraries from the same RNA tube are technical replicates; they estimate machine noise,not biology, and should be summed rather than treated as separate samples. Power in RNA-seqcomes far more from adding biological replicates than from adding depth. Three replicatesper group is the prac- tical floor, five or six is comfortable.A batch is any grouping of samples that share a processing nuisance: extraction day,reagent lot, flow cell, technician. A variable is confounded with the biological factorwhen knowing one tells you the other. Randomisation spreads unknown nui- sances evenly;blocking takes a known nuisance and forces every level of the biological factor to appearinside every block, so the nuisance can be estimated and removed in the model. Confounded batch A batch B C1 C2 C3 T1 T2 T3 batch effect and treatment effect cannot be separated Balanced and blocked batch A batch B C1 T1 C2 T2 C3 T3 C4 T4 batch enters the design as a term, treatment stays estimable Figure 2: Left, treatment is perfectly confounded with batch. Right,