Bioinformatics Notes

Single Cell Genomics and Multi Omics Integration

Lecture slides and notes for Single Cell Genomics and Multi Omics Integration in Bioinformatics Notes by Md Ahbab. 13 pages.

Document Info: 13 pages · PDF

Single Cell Genomics and Multi Omics Integration, first page preview

Content Preview

CSE 4893: Introduction to Bioinformatics Module 12: Single Cell Genomics and Multi Omics Integration Abstract Single cell genomics measures the transcriptome of individual cells, so populations that a bulkexperiment averages away become visible. These notes follow one dataset from raw reads to annotatedcell types. We cover barcoding chemistries, count matrix construction with Cell- Ranger, STARsolo andalevin-fry, quality control for empty droplets, ambient RNA and doublets, normalisation, featureselection, principal component analysis, graph clustering and two dimen- sional embeddings. We thentreat batch integration, pseudobulk differential testing, trajectory inference, RNA velocity, spatialplatforms and factor models for multi omics. Worked exercises, algorithms and a keyword glossarysupport revision. Throughout, the emphasis is on which choices are biology and which are ours. Contents 1 Learning Objectives 2 2 Why Single Cell 2 3 Technologies 2 4 From Reads to a Matrix 3 5 Quality Control 3 6 Normalisation and Feature Selection 4 7 Dimensionality Reduction and Clustering 4 8 Embeddings 5 9 Annotation and Differential Testing 5 10 Batch Integration 6 11 Dynamics 6 12 Spatial and

CSE 4893 Introduction to Bioinformatics Module 12 19 Useful Links 12 1 Learning Objectives 1. Explain why bulk averaging hides opposing cell populations, and judge when single cell resolution is the right instrument for a biological question. 2. Describe how sequencing reads become a sparse count matrix, including the roles of the cell barcode and the unique molecular identifier. 3. Apply quality control, normalisation, feature selection and dimensionality reduction with thresh- olds you can defend from the data rather than from habit. 4. Compare batch integration methods and reason about the trade off between mixing batches and conserving biological variation. 5. Interpret trajectories, RNA velocity, spatial assays and multi omics factor models together with the assumptions that make each of them work. 2 Why Single Cell A bulk RNA sequencing experiment reports one number per gene for a whole tissue. That number is aweighted mean over the cells present, xbulk g = P c πcxgc, where πc is thefraction of the library contributed by cell c. The mean is honest but it is not informative when twopopulations move in opposite directions. A gene that doubles in fibroblasts and halves in macrop

CSE 4893 Introduction to Bioinformatics Module 12 Table 1: Representative properties of the three assay families. Costs are order of magnitude reagentcosts per cell and exclude sequencing and labour. Method Cells per run Reads per cell Coverage Cost per cell Plate based (Smart-seq2, Smart-seq3) 100 to 1000 500 000 to 1 000 000 full length, isoforms and variants $5 to $20 Droplet based (Chromium 3’ and 5’) 500 to 20 000 per lane 20 000 to 50 000 3’ or 5’ end tag only $0.10 to $0.50 Combinatorial indexing (sci-RNA-seq3, SPLiT-seq) 100 000 to 1 000 000 1000 to 10 000 end tag, often nuclei only below $0.05 4 From Reads to a Matrix Quantification takes FASTQ files and returns a genes by cells matrix. CellRanger (version 10.1,2026) is the vendor pipeline for Chromium data and now ships local cell type annotation. STARsolo , built into the STAR aligner, reproduces CellRanger output with a genome alignment andis typically several times faster. alevin-fry, driven by the simpleaf wrapper, usesselective alignment against a spliced transcriptome and is the lightest of the three in memory and disk.All three perform the same four logical steps shown in Figure 2: barcode correctionagainst a known